Three services, each scoped in writing and priced before it starts. Every engagement ships with the parts teams usually defer: tests, observability, runbooks, and documentation your next engineer can follow.
Most products don't fail on features — they fail the first time load doubles. We build the layer underneath: well-shaped APIs, idempotent jobs, sane queueing, and observability you can actually debug at 3am.
We work in your stack rather than imposing ours. TypeScript, Go, Python, Postgres, Redis — chosen because your team can maintain it after we leave, not because it's fashionable this quarter.
deliverables
Architecture decision records for every significant choice
Versioned REST or GraphQL API with generated docs and typed clients
Background job and queue infrastructure with retry semantics
Structured logging, tracing, and alerting wired to real thresholds
Load-test results against agreed targets, not vibes
Runbooks for the five failures most likely to wake someone up
An LLM demo takes an afternoon. Making it reliable enough to put in front of customers takes engineering: retrieval that returns the right chunk, evaluation that catches regressions, and guardrails that fail closed.
Every engagement starts with an honest assessment of whether the model can hit the bar your use case needs. When it can't, we tell you before you've spent the budget rather than after.
deliverables
Feasibility assessment with a documented go / no-go recommendation
Ingestion and transformation pipelines with data-quality checks
Retrieval-augmented generation over your own corpus
Evaluation harness with regression suite and scored baselines
Prompt and model versioning, plus rollback capability
Token and inference cost model at your projected volume
Claude APIOpenAIpgvectorLangGraphdbtAirflowPythonDuckDB
typical_shape
Fixed scope, fixed price, milestone billing
Two-week cycles ending in a deployed environment
Shared channel with the engineers writing the code
Legacy systems are rarely as broken as they feel — they are usually undocumented, untested, and frightening to touch. Those are different problems with different, cheaper fixes.
We start read-only: map the system, instrument it, and find where the risk actually concentrates. Only then do we recommend what to change. A full rewrite is the answer roughly one time in ten, and we will say so when it isn't.
deliverables
Dependency and data-flow map of the existing system
Prioritised risk register with severity and blast radius
Characterisation tests around the parts you're afraid to touch
Incremental migration plan with reversible steps
CI pipeline and environment parity where none existed
If an existing product solves it, we name the product
If the budget cannot carry the scope, you hear it in call one
If a model cannot hit your reliability bar, we will not ship it
If a rewrite is not warranted, we argue against our own upside
Describe the problem, not the solution.
Tell us what is breaking, what is slow, or what will not scale. We will tell you what it takes to fix — including when the answer is that you should not build it.