Technical architecture
{% raw %}
This document describes the proposed technology choices for Crabster and the code it generates. Each major choice is tagged [ADR] and can be revised — but any revision must be documented as such (status, alternatives considered, reason for the change).
1. Overview🔗
Crabster is made of two distinct parts:
2. The generator (CLI)🔗
2.1 Language and distribution — [ADR]🔗
- Rust, 2021 edition (2024 migration possible once stabilized across the dependencies used).
- CLI built with
clap(derive API):crabster new— initializes a new project. Shipped.crabster import-cdl <file>— generates from a full.cdlfile. Shipped.crabster record <name> --field "..."— adds a record to an existing project. Shipped.crabster upgrade— updates a generated project to a newer template version (the trickiest part, see §6). Planned, Phase 11.
- Distribution:
cargo install crabster-cliin V1; prebuilt binaries (viacargo-distorcargo-binstall) once the tool stabilizes, to avoid requiring a full Rust toolchain for users who would only be generating, say, a Node/Java project (not applicable to the API-only V1, but relevant if scope widens).
2.2 CDL parsing🔗
-
Formally defined grammar (see CDL language), parser written with
pest— [ADR settled in Phase 3, 2026-09-05].Chosen over
chumskyfor three reasons. The grammar is a specification artefact that gets reviewed: a separate.pestfile, close to BNF, serves that purpose where combinators buried in Rust would not.pestreports line, column and expected tokens out of the box, which Phase 3's Definition of Done requires. And CDL is a declarative language with no expressions: no precedence, no ambiguity, so none of the hard cases where chumsky's error recovery earns its keep.The two were not prototyped as the roadmap allowed for: the doubt did not survive the comparison, and building both would not have paid for itself. Should pest's diagnostics prove too thin in practice, the parser sits behind the IR and can be replaced without touching a single template.
-
The parser produces an Intermediate Representation (IR): Rust structures representing records, fields, types, references, validations, generation options — independent from CDL syntax itself, to allow adding other model sources later (importing an existing Swagger/OpenAPI spec, reflecting on an existing SQL schema, etc.).
2.3 Template engine — [ADR]🔗
-
Tera(Jinja2/Django-like syntax) rather thanaskama: Tera accepts templates supplied at runtime, which is essential for the blueprint mechanism (customizing/overriding without recompiling the generator).askamacompiles templates into the binary (faster, more type-safe) but would make external blueprints impossible without redistributing a custom binary — incompatible with the community extensibility goal. -
Refined in Phase 2: templates do not all come from disk, deliberately. A CLI installed through
cargo installhas notemplates/directory beside it — a disk-only engine would ship a generator that cannot generate. Built-in modules are therefore embedded in the binary (include_dir!) and handed to Tera throughadd_raw_template, while blueprints are read from disk. The ADR's goal is intact — a blueprint author drops in a directory, without rebuilding Crabster — and the binary stays self-contained.This has a consequence for the repository layout: the template tree lives under
crates/crabster-codegen/, not at the workspace root.cargo packageonly ships what sits below the package root; a directory beside it would not enter the.crate, andinclude_dir!on a missing directory is a compile error, not a silent degradation.cargo install crabster-cliwould therefore fail for every user. ThepackageableCI job packages all three crates and rebuilds each from its own tarball, which makes that defect impossible to reintroduce. -
Template layout: one directory per generatable "module" (
core,record,auth-jwt, …), each with its.terafiles and amodule.toml(inter-module dependencies, expected context variables, what turns the module on, the names it occupies). A template's destination is derived from its path by dropping.terarather than declared — one source of truth, so it cannot drift. -
A blueprint is a directory of the same modules, stacked on top. Resolution is per template, not per module, so a blueprint that provides one file of
coreinherits every other. A template that renders only whitespace writes no file, which is how a blueprint removes one. See Writing a blueprint. -
A blueprint module declares the templates version it was written against, and is refused otherwise. It writes against slots, variables and file names that move — before 1.0 a minor version may change any of them — and without the declaration the failure arrives from inside a template, saying nothing about the cause. The built-in modules do not declare it: they ship in the same binary as the number they would be compared against.
-
Generated Rust is auto-formatted with
rustfmt(best-effort: a missingrustfmtdoes not fail generation), and can be checked withcargo checkthroughcrabster new --check. That check is opt-in rather than automatic: it compiles the whole dependency tree, turning an instant command into a minute-long one.
2.4 Blueprint mechanism (extensibility)🔗
The mechanism:
- A blueprint is an external crate or directory providing
.teratemplates that replace or extend the core ones. - Layered resolution:
user blueprint > installed community blueprint > core templates. Overriding happens per template, not per module: a blueprint providingcore/README.md.terareplaces that one file and inherits the rest, instead of forcing its author to vendor and maintain a whole module to change one line. - A community blueprint registry (eventually, a simple site listing crates.io crates tagged
crabster-blueprint) enables discovery. - A template that renders to nothing but whitespace writes no file. That is how a module says "not this time" —
src/domain/enums.rs.terawraps itself in an{% if model.enums | length > 0 %}, so a model with no enumeration does not get a file nothing declares or reads. No Rust-side registration is involved, which is the point; the price is that a deliberately empty file cannot be generated this way. - Two guards apply to blueprints, which are third-party code: symlinks are skipped rather than followed — including a module directory that is itself a link, which then makes the blueprint report that it holds no module rather than silently falling back — and a destination escaping the generated project is refused.
3. The generated project (backend)🔗
3.1 Web framework — [ADR]🔗
- Axum, chosen over Actix-web or Rocket for:
- Native integration with the
tower/tower-httpecosystem (reusable middleware: compression, CORS, timeouts, tracing, rate-limiting). - Maintained by the Tokio team, wide and growing adoption, good compatibility with
sqlx/sea-orm. - No Rocket-style macro magic; explicit types, easier for generated code to be read and understood by a developer new to the codebase.
- Native integration with the
- Documented but not chosen for V1: Actix-web (comparable performance, mature ecosystem, but its actor model and macros are less aligned with a "readable generated code" style).
3.2 ORM and data access — [ADR]🔗
- SeaORM, chosen over Diesel for:
- Native async API (consistent with Axum/Tokio), whereas Diesel is synchronous by default (requires a blocking pool +
spawn_blocking, extra complexity to generate and explain). - An entity/relationship model familiar to anyone who has used an ORM (repositories,
has_many/belongs_torelations, configurable loading). - Built-in migrations (
sea-orm-migration), code generation from an existing schema (sea-orm-cli generate entity) useful for a future "reverse engineering" mode (see Phase 13). - Diesel remains documented as a possible alternative blueprint for teams preferring compile-time query checking.
- Native async API (consistent with Axum/Tokio), whereas Diesel is synchronous by default (requires a blocking pool +
- V1 database support: PostgreSQL (reference), MySQL (written
mysql; MariaDB is wire-compatible), SQLite (dev/tests). NoSQL (MongoDB) documented as a future extension outside SeaORM (dedicated driver).
3.3 Migrations🔗
sea-orm-migration, migration files generated and versioned undersrc/migration/, as a Rust module of the crate — one file per record, numbered in the order they must run, native Rust and compiled: no separate XML or YAML alongside the code.
3.4 Authentication and security — [ADR]🔗
- Stateless JWT via
jsonwebtoken, generated by theauth-jwtmodule, whichservice { auth jwt }turns on.rust_cryptorather thanaws_lc_rs, so a generated project needs no C toolchain to build — and one of the two is required, since without either the crate panics at the first signature, at runtime. - Password hashing:
argon2, at the crate's recommended parameters rather than pinned ones — the encoding stores what it was made with, so a hash written under the old parameters still verifies after an upgrade. A login that finds no account hashes against a dummy anyway, so a missing address is not measurably faster than a wrong password. - Authorization: an
Authenticatedextractor, andclaims.require_role("ADMIN"). Not a macro, which is what this section used to promise: a method reads the same as the code around it, appears in a stack trace, and needs no explanation of what it expands to. Naming the extractor in a handler is what makes a route protected, so a route cannot be listed as protected and not be. - Roles are read from the token, and re-read from the account on every refresh. That is what stateless costs: a role taken away takes effect when the access token expires — fifteen minutes — not when it is taken away. A token cannot be revoked at all, which is why the access lifetime is short and the refresh lifetime is the long one.
- The signing key is refused twice: the placeholder
.env.exampleships is refused outside thedevprofile, and a key shorter than 32 bytes is refused in every profile. Both at startup, so a service that cannot sign anything never gets as far as being called healthy. - HTTP limits are
core's, not the auth module's: a request timeout answered as408, a concurrency limit, a body limit,nosniffon every response, and CORS that is closed by default with no way to say "any" — a wildcard on an API that reads a token is a wildcard on that token. - No rate limiting is generated, deliberately. A limiter inside the process counts one replica's traffic, so the configured number means something different every time the deployment is scaled, and what is worth limiting is usually per caller — an identity that layer does not have. It belongs at whatever terminates TLS in front of the service.
- OAuth2/OIDC: deferred past V1, via the
oauth2crate — documented as a later phase rather than inflating V1 scope. - Session-based auth: not a priority, since stateless JWT covers the most common API-only use case.
3.5 API documentation — [ADR]🔗
utoipa+utoipa-swagger-ui: generates the OpenAPI 3.1 document directly from attributes on the generated handlers and DTOs ("code-first", consistent with Rust style, rather than a separate OpenAPI file to keep in sync by hand). Served at/api-docs/openapi.json, with a console at/swagger-ui.- The Swagger UI assets are vendored, not fetched.
utoipa-swagger-uidownloads them in its build script by default, which would make a generated project unbuildable without a network and pullreqwestinto its build tree; thevendoredfeature compiles them in instead, at a cost of about two megabytes. - Every operation carries an explicit
operation_id. A generator produces alistin every record's module, and a specification may not name two operations the same — so the identifier is built from the record's name rather than left to default to the function's. - The types SeaORM aliases are annotated explicitly. utoipa reads a type by the name it is written as, and
Date,Decimal,UuidandDateTimeUtcare SeaORM's aliases — names it has never heard of. The mapping to a schema type and format lives inview.rs, beside the Rust type and the column type, which are the other two halves of the same decision. - A
.spectral.yamlis generated with the project, and the document passes it with nothing at any severity. One rule is turned off,info-contact: who answers for an API is decided by whoever runs it, not by whoever generated it.
3.6 Validation🔗
validatoron generated DTOs, rules auto-derived from CDL constraints (required,minlength,maxlength,pattern,min/max).
3.7 Configuration🔗
config+ environment variables, profiles named byAPP_PROFILE(config/{profile}.toml). Two are generated:prod, which is what an absentAPP_PROFILEgets — a deployment has no.env, and a server binding the loopback address by default would be a container nobody could reach — anddev, which.env.exampleselects. Noconfig/test.tomlis generated: the generated tests build the app in-process against an in-memory SQLite and never readconfig/. Anyconfig/{name}.tomldropped there is picked up.- Secrets never committed:
.env.examplegenerated, documented integration with external secret managers (Vault, AWS Secrets Manager…) as a guide, not an imposed dependency.
3.8 Errors🔗
- A unified application error type via
thiserrorfor internal errors, converted to a structured HTTP response through itsIntoResponseimplementation. - Every failure answers RFC 9457 problem details — the revision of RFC 7807 — as
application/problem+json, including the ones axum would otherwise reject itself with plain text or an empty body. typeis what a client branches on, and the field the status cannot replace: a409for a taken unique value and a409for a reference that does not hold are different problems under one code. It is a relative URI the project serves:GET /problems/unique-conflictexplains that kind andGET /problemslists them all. RFC 9457 says a type URI that is a locator should have documentation behind it, so it does — rather thanabout:blank, which would claim there is nothing to say.- A refused body is 422, not 400, and the roadmap's Definition of Done was corrected rather than the code: RFC 9110 defines 422 as a request that is well formed and semantically wrong, axum already answers 422 one layer out for a body that does not match the shape, and collapsing both into 400 would make "I could not read this" indistinguishable from "I read it and will not have it".
3.9 Tests🔗
- Standard Rust unit tests (
#[cfg(test)]), driving the whole router rather than calling handlers — routing, extractors, layers and the error shape are exactly what a test that calls a handler directly does not exercise. - Against the database the project targets, via
testcontainers:src/testing.rsstarts one container for the whole test binary and creates a database per test. SQLite needs neither, being a file.TEST_DATABASE_URLbypasses the container for a CI job that already runs a server. - Against the search engine it targets, the same way:
src/search/testing.rsstarts one Meilisearch for the whole test binary and hands each state a set of index names nobody else is using.TEST_SEARCH_ADDRESSbypasses the container. A suite that expected somebody to have started an engine by hand would pass on one machine and fail on every other, which is not a suite. - And it takes them away again. A handle in a
staticis never dropped, so nothing givestestcontainersthe chance to remove what it started. Each harness spawns a reaper holding the read end of a pipe the suite never writes to: when the run ends, however it ends, the kernel closes the pipe and the reaper removes the container. A destructor would have missedSIGKILL; this does not. - A database per test, not a shared one. These tests count rows; sharing one would make them depend on the order they happened to run in.
- Every record is covered, including the two shapes that used to be skipped. A mandatory reference is created through the parent's own endpoint, and a
@matchessample is generated from the pattern and then checked against it — a table of well-known patterns would cover the three everybody writes and leave the rest of the language untested. - A full CRUD test suite generated per record (create/read/update/delete/list/filter/pagination).
3.10 Observability🔗
tracing+tracing-subscriberfor structured logs — readable in thedevprofile, JSON everywhere else. OpenTelemetry export stays optional and is not generated: what a service ships its traces to is decided where it is deployed, not in its code. The hook is in place (init_tracinginsrc/observability.rs) and the generated README describes the change.- Correlation identifier. Every request carries an
x-request-id, minted if it arrived without one, written into every log line of that request, and returned on the response. - Prometheus metrics endpoint via
axum-prometheus: a counter, a latency histogram and in-flight requests, labelled by the matched route —/api/products/{id}, never/api/products/1— so the number of series is bounded by the size of the project rather than by its traffic. The health probes and/metricsitself are not counted. - Three health endpoints.
/health/livereaches nothing: a restart policy reads it, and restarting a healthy process because a database went down turns one outage into two./health/readyreaches everything the service depends on and answers503naming what is missing; a load balancer reads it./healthremains, answering liveness, under the name it had before the two were split apart. - Readiness is composed, not hardcoded.
src/health.rsbelongs tocore, which knows nothing about a database; the probe that pings one is contributed by therecordmodule through the slots in that file. A project with no domain model keeps the endpoint and has nothing to list — and a module that later brings a dependency of its own adds its probe withoutcorechanging.
3.11 Containerization and deployment🔗
- Generated multi-stage
Dockerfile, with dependency caching viacargo-chef—cargorebuilds everything when any file changes, so without it a one-line edit recompiles the whole tree. The final stage isgcr.io/distroless/cc: the binary, itsconfig/, and nothing else, running asnonroot.ccrather thanstatic, because the build links against glibc. - The base image is pinned to the project's own MSRV, which lives in each module's manifest and reaches the templates as one variable. Three files name that number — the manifest, the
Dockerfile, and the CI the README suggests — and three copies is three chances to raise two of them; a test holds the first two together. - Generated
docker-compose.ymlfor local development. The database has been there since Phase 5, on exactly the host, port, user, password and database name.env.examplepoints at, socp .env.example .env && docker compose up -d --wait && cargo runneeds nothing filled in. Nothing is generated for SQLite, which is a file and has no server to start. The app joins it in Phase 8, along with theDockerfileit needs to build. Any further service goes in adocker-compose.override.yml, which Compose merges in on its own and which regenerating never touches. - Optional base Kubernetes manifests generated (Deployment, Service, ConfigMap) — low priority, Phase 11.
3.12 CI/CD🔗
- Generated GitHub Actions pipeline:
cargo fmt --check,cargo clippy -- -D warnings,cargo testagainst the database the project targets,rustsec/audit-check, and a build of the image. No secret and no repository setting, so it passes on a repository that has just been created. - The image is built and not pushed. Where an image belongs is a decision about infrastructure, and pushing needs a registry and a credential a generator cannot invent; the workflow says in a comment what to add.
- GitLab CI is documented rather than generated — you use one or the other, and generating both leaves a dead file in every project. The generated README carries the equivalent pipeline in full.
4. Intermediate Representation (IR) — the central pivot🔗
The IR is the stable contract between "CDL parser" and "generation engine". It must stay agnostic of CDL syntax to allow, eventually, other input sources (importing an existing SQL schema, importing an OpenAPI spec). Sketch of the structure (detailed in Phase 3):
struct DomainModel {
records: Vec<Record>,
enums: Vec<Enumeration>,
service: Option<Service>, // name, database, port
}
struct Record {
name: String,
fields: Vec<Field>, // a field may be a reference
filterable: bool,
}
struct Field {
name: String,
kind: FieldType, // Scalar | Enum(name) | Reference(name)
optional: bool, // mandatory by default
unique: bool,
length: Option<Bounds>,
range: Option<Bounds>,
matches: Option<String>,
}
There is no relationship type: a reference is a field, and the side declaring it is the side holding the foreign key. That absence removes all the code that would otherwise decide which side the column belongs on.
5. Crabster repository layout (proposed monorepo)🔗
crabster/
├── crates/
│ ├── crabster-cli/ # binary, clap subcommands
│ ├── crabster-cdl/ # CDL parser + IR
│ ├── crabster-codegen/ # Tera engine + module orchestration
│ │ └── templates/ # templates for the generated code — under the crate,
│ │ ├── core/ # not at the root: `cargo package` only ships what is
│ │ ├── record/ # below the package root, and `include_dir!` on a
│ │ ├── auth-jwt/ # missing directory is a compile error
│ │ ├── db-postgres/ db-mysql/ db-sqlite/
│ │ ├── docker/
│ │ └── ci-github-actions/
│ └── crabster-shared/ # shared utilities (if needed at generated-app runtime)
├── examples/ # reference CDL models, generated and exercised in CI
└── docs/
6. The incremental-update challenge ("upgrade")🔗
The hardest problem in this space: merging generated code the user has since modified. Planned approach:
-
V1: "one-shot" generation for a new project, plus simple record addition (
crabster record). The addition merges nothing: every generated file was stamped as it was written, and a file still matching its stamp is ours and may move. One that no longer matches was edited by its owner, and the command stops and names it, having written nothing, unless--forcesays otherwise.A project therefore keeps, in
.crabster/:model.cdl, the original.cdlverbatim — comments included;project.toml, what the command line decided and no.cdlcan express (--database,--port, the name); andfiles.toml, one stamp per generated file. The obvious alternative — render the old model again and compare — looks cheaper and does not work: rendering runs throughrustfmt, an external binary found onPATH, unpinned, and answering to anyrustfmt.tomlin the working directory. A toolchain installed withoutrustfmtmade eighteen untouched files look hand-edited. A stamp taken at the moment of writing depends on neither. It also keepssnapshot/, a copy of every generated file as it was written: that copy is the common ancestor of the three-way mergecrabster upgradeperforms on a file its owner edited, and there is no other way to have one — what the generator produced on a given day cannot be recovered afterwards, its templates being compiled into the binary that wrote it. It costs the repository one committed copy of its generated code. The directory is committed with the project, and it is the.crabster/ADR-0002 describes.Three guards apply. A project whose files no longer match what it records — a generated file deleted, a model edited by hand — is refused, since otherwise migrations already applied would be renumbered and orphan modules left behind. A record may only be appended, which leaves existing migration names untouched. And a project generated by a different version of Crabster is refused too: rendering its model again says what the files held only while the templates have not moved, and otherwise the whole project would look hand-edited. Crossing a version is
crabster upgrade, in Phase 11. No smart merging of hand-modified files either way: Phase 11 as well. -
Later phase: generated files marked with protected-zone comments (reserved-zone markers), enabling assisted (three-way,
git merge-style) merging when template versions are upgraded.
7. Alternatives explicitly ruled out (and why)🔗
| Chosen | Alternative ruled out | Reason |
|---|---|---|
| Axum | Rocket | Macros too "magic" for generated code meant to be easily read/modified |
| Axum | Actix-web | Actor model adds unnecessary conceptual complexity; Axum/tower is sufficient |
| SeaORM | Diesel | Native async API, familiar entity/relationship model |
| SeaORM | Raw SQLx | SQLx remains a viable option for a "hand-written SQL" blueprint, but SeaORM better matches the default "entities" experience expected |
| Tera | Askama | Runtime-loaded templates are required for external blueprints |
| Stateless JWT | Session | API-only use case is the priority; session support deferred |
| {% endraw %} |