Agentic Systems Engineering·Claude
The deliverable is a running system. We stay on the hook for it.
We design and build agent systems on Claude — from whether the problem is tractable at all, through architecture and the build, to production and the evaluation that keeps it honest.
Forward deployed means our engineers work inside your team, in your codebase, through your review process — and are accountable for whether the system still runs a year from now.
We lead with Claude because that is where our customers have already invested. The harness itself stays model-neutral, because measurement tied to a single vendor is not measurement — you need a fixed yardstick to tell whether a change came from the model, your prompts, or your data.
How an engagement runs
Five stages. You can start at any of them — most teams come to us somewhere in the middle — but a practice that can only do the last one is not much use at the first.
- 01
Assess
Is this tractable with agents at all? We look at the problem, the data it depends on, where it will fail, and what it costs to run.
You keep
- A feasibility judgment, including when the answer is no
- A candidate architecture and its failure modes
- Operating cost and latency you can plan against
- 02
Design
Task decomposition, orchestration topology, tool boundaries, state and memory, where a human stays in the loop, and what happens when a step fails.
You keep
- A system design validated by building the risky part
- Tool and MCP interfaces with explicit permission boundaries
- Recovery and escalation behavior, specified
- 03
Engineer
The build. Agent SDK and orchestration, MCP servers, Claude Code in your workflow, context and cost engineering — as production code in your repo, through your review process.
You keep
- Code your engineers reviewed and can maintain
- Deterministic guardrails, enforced in hooks rather than prompts
- Cost engineering — caching, batching, model selection
- 04
Productionize
Deployment, observability, cost and latency profiling, security and compliance review. We hand over rather than stay forever.
You keep
- Runbooks and an on-call handover your team can hold
- Observability wired to your existing stack
- Security and compliance review passed, not deferred
- 05
Evaluate
Eval suites in CI that block merge on regression, drift detection, and automated re-baselining when a provider ships a new model version.
You keep
- Automated re-baselining on every model release
- Regression, latency and cost deltas before you upgrade
- Evidence packages your auditors accept
What we build
- Agents & MCP servers
- Production agents and MCP servers: orchestration and subagents, tool interfaces with explicit permission boundaries, state and recovery behavior, observability, runbooks. Built to pass your security review, not to demo.
- Claude Code enablement
- Rolling Claude Code across an engineering organization: CLAUDE.md and rule hierarchy, hooks that block destructive operations deterministically, skills and subagents, CI integration, and measurement of whether it actually helped.
- Context & cost engineering
- The two things that quietly sink production Claude systems. Context compaction and long-horizon state, prompt caching, batching, and model selection — tuned against what the workload actually costs to run.
- Evaluation harness & model-upgrade regression
- Production eval suites in CI that block merge on regression, plus automated re-baselining when a provider ships a new version. Almost nobody sells this as a service. It is the work we are best at.
- Code modernization
- Legacy migration and test-suite reconstruction using agentic coding. Characterization tests first, so the refactor is pinned by behavior rather than hope.
- Embedded pod
- A standing team inside your engineering org. Available after a delivered engagement — never as a starting point, because an embedded team sold cold is just contract staffing.
Eight to sixteen weeks
Four weeks
Three to six weeks
Six to ten weeks
Twelve weeks and up
Quarterly
How we build
These are not our opinions. They are Anthropic’s published engineering guidance, and they are what we hold ourselves to on review.
- The simplest thing that works
- Anthropic’s own guidance is to prefer a workflow over an agent, and add agentic complexity only when a workflow genuinely cannot do the job. A firm that sells agents has every incentive to ignore that. We do not get to.
- Guardrails in code, not in prompts
- A prompt is probabilistic. Business-critical rules belong in hooks and programmatic prerequisites, where compliance is guaranteed rather than likely.
- Fewer tools, better described
- Agents degrade as tool counts climb. We consolidate rather than wrap every endpoint, and write tool descriptions as carefully as the tools themselves.
- Nothing reviews its own work
- Self-review inside the session that produced the output catches far less than an independent pass. That applies to the agents we build and to us.
What we commit to
A system running in production, not a set of recommendations. Every engagement has a defined end and a handover — your engineers own what we build and can change it without us.
We will tell you early if we think the answer is no. A workflow that a script can do reliably should not become an agent, and you should hear that from us before you spend on it.
How an engagement is structured — the five phases, where we interlock with your teams, and how it is contracted — is set out on the engagement model.
Capacity and terms are sized per engagement. Ask and we will walk you through it.
Who this is for
- You are automating enterprise workflows and want to know where agents genuinely fit — and where they do not.
- You have not started yet, or you started and stalled short of production.
- You are not yet confident which workflows are worth the investment.
- Someone senior owns the outcome and can clear a path.
- You carry real obligations — SOC 2, model risk, sector regulation.
- Your engineers will work alongside ours, not hand us a specification.
Not knowing where to start is a normal place to be, and it is a good place to bring us in — the first stage exists precisely to answer it. What you get out of it is a judgment you can act on, including the judgment that a particular workflow is not worth building.
Tell us what is stuck.
The first conversation is with the engineers who would do the work, not a sales qualification.