The engagement

Five tiers, from one developer to every team.

Each tier ships something your engineers keep. Open "trust controls" to see how identity, guardrails, evaluations, observability and cost are handled at that tier.

1

Day one · The individual

Orientation

A safe, accountable way for every individual to start using AI.

We shipBusiness AI access with everyone signed in as themselves, a one-page "how we use AI" note, and a first labelled, human-reviewed pull request.

Trust controls
Identity
A business or team AI plan, with everyone signing in as themselves. No personal accounts or shared logins.
Safety & Guardrails
Training on your data is turned off, no secrets or customer data go into prompts, and a human reviews all AI-written code.
Evaluations
People check outputs against work they know well.
Observability
AI-assisted pull requests are labelled, and useful prompts are kept.
FinOps
Fixed seats, with one person checking the bill monthly.
2

Day two · The team

Implementation

Move from individual use to shared agents and skills, so the whole team is enabled and work is transparent.

We shipShared agents and skills in the repository with named owners, team rules files (for example AGENTS.md), and a shared list of real tasks with known good results.

Trust controls
Identity
Shared agents and skills live in the repository with named owners and run with the user's own permissions.
Safety & Guardrails
Team rules files set standards and no-go areas. Pull request review and branch protection apply to AI changes just as they do to human changes.
Evaluations
Reference tasks are re-run whenever a shared agent, skill or model changes.
Observability
Usage shows who uses which agents and skills; artifacts trace to the skill or prompt that produced them, and everything is version-controlled.
FinOps
A team spend cap, with usage reviewed by person and by tool.
3

Two-week engagement · Reusable components, templates and skills

Distillation

Agents assemble proven parts instead of generating from scratch: faster delivery, fewer tokens, more consistent output.

We shipVersioned components, templates and skills that encode your architecture, security and coding standards, with reference tasks and a before-and-after baseline.

Trust controls
Identity
Every component, template and skill has an owner and a version, so every artifact has a clear source.
Safety & Guardrails
Templates build in your standards, so the safe path is the default. Less for the AI to invent means less to go wrong.
Evaluations
Each skill has reference tasks with known good results; before-and-after comparisons show quality holding while effort drops.
Observability
Usage shows which skills and templates are used, skipped or hand-edited; artifacts trace to the component and template that generated them.
FinOps
Tokens and time per task, measured before and after: your proof of value and baseline.
4

A few weeks · The delivery pipeline

Expanding

Extend trusted AI use into the delivery pipeline.

We shipAgents in CI or the cloud with scoped service identities, an allowlist of tools and MCP servers, reference tasks as CI quality gates, and spend alerts.

Trust controls
Identity
Agents get their own service identities with narrowly scoped tokens; humans remain the approvers.
Safety & Guardrails
An allowlist of tools and connectors (MCP servers), secret scanning and prompt-injection checks.
Evaluations
Tier 3 reference tasks become CI quality gates and regression suites.
Observability
Agent runs log the components, skills and tool calls used, linked to the pull request or ticket.
FinOps
Costs tagged per project or pipeline, with alerts for unusual spend or runaway agent loops.
5

A few months · Across teams

Scaling

Scale trusted AI use across teams and the wider organisation.

We shipA central register of agents, skills and components, platform-enforced policies, a shared evaluation library, and cost-per-feature reporting.

Trust controls
Identity
A central register with owners, permissions and lifecycle, tied to company login and access reviews.
Safety & Guardrails
Platform-enforced policies, red-teaming, and alignment with ISO/IEC 42001 or the NIST AI RMF.
Evaluations
A shared evaluation library, with production quality monitored against it.
Observability
An organisation-wide view of reusable assets and what they produce, using OpenTelemetry GenAI conventions.
FinOps
Cost per task or feature, model routing by cost versus quality, and chargeback against the Tier 3 baseline.

Why Context Compression matters.

When agents assemble approved parts rather than inventing from scratch, guardrails, evaluations, observability and cost control all become easier at once.

It's the bridge from team use to the pipeline: the components, reference tasks and baselines created here are what Tiers 4 and 5 rely on.

Value evidence

How you'll know it's working.

Measured against the baseline set in Tier 3, so the numbers stand up to scrutiny. How we make ROI trustworthy

Cycle time

Time from task start to merged, reviewed change.

Acceptance and rework

Share of AI-assisted changes accepted without significant rework.

Escaped defects

Quality in production, not just speed of delivery.

Cost per task or feature

Tokens, time and tooling against delivered work.