Insights

Best practices for adopting agentic AI in the SDLC with Anthropic’s Claude


An enterprise adoption guide to context engineering, harness design, oversight, and cost governance.

Adnan Masood, PhD, Chief AI architect, UST.

Most organizations have adopted AI coding tools. Few have redesigned software delivery around them. Learn how to build the governance, harness, and oversight model required to achieve measurable engineering outcomes.

Adnan Masood, PhD, Chief AI architect, UST.

The productivity plateau is real

Gartner reports that about 27% of engineering organizations now use agentic AI somewhere in the software development life cycle. Yet roughly 71% of software engineering leaders report gains of only 1% to 25% from AI tools . That is a respectable pilot result and a disappointing transformation result.

The pattern behind the plateau is familiar to anyone who has run a delivery organization. Coding gets faster, so the queue moves to code review. Review gets faster, so the queue moves to testing. Engineers spend the saved hours supervising, validating, and reconciling agent output. Throughput stays flat, and the finance team starts asking pointed questions about the token bill.

Gartner’s central claim deserves attention: the agent harness determines output quality more than the choice of model . The harness is the layer that feeds the agent context, limits what it can touch, and checks its work before anyone accepts it. With Claude, that harness is Claude Code, running interactively, headless, or in CI, or a custom agent built on the Claude Agent SDK. Every part of it is a file you can version, review, and ship.

DIVIDER

1. Start from the constraint and measure flow

Point agents at the bottleneck that limits end-to-end flow. If pull requests (PRs) wait 26 hours for code review, a faster code generator makes that bottleneck worse.

Instrument before you pilot. Claude Code exports OpenTelemetry metrics for sessions, tokens, cost, lines changed, commits, and pull requests. Team and Enterprise plans add an analytics dashboard with PR attribution. Establish the baseline and the treatment in the same pipeline, and report cost per merged PR by use case. Leadership understands that unit.

Then choose the intervention per use case. Formatting and dependency bumps belong in deterministic automation. Design exploration and debugging suit an interactive assistant. Test backfill, flaky-test triage, and migration fan-out are agent work, provided the repo has the test coverage to catch mistakes.

DIVIDER

2. Ship a paved road for agents

Adoption fails when every team builds its own configuration. Claude Code supports private plugin marketplaces that bundle skills, subagents, hooks, and MCP servers, and managed settings that require them across the fleet . Treat that bundle as the internal developer platform for agents. Teams install one vetted baseline instead of copying snippets from an outdated Confluence page.

DIVIDER

3. Write down what your senior engineers know

Agents arrive with broad knowledge and zero memory of your architecture. Context engineering fixes this with version-controlled artifacts. In Claude Code, a root CLAUDE.md holds build commands, architectural invariants, and the definition of done. Path-scoped rules in .claude/rules/ load only when Claude touches matching files, so migration conventions appear only during migration work. Skills package reusable procedures such as writing an ADR, and MCP servers supply live context from the service catalog and issue tracker .

Treat this context like code. Put it under CODEOWNERS, and schedule a headless Claude run that compares CLAUDE.md against the actual build and recent merged PRs and opens a PR for anything stale. Context rot is quiet, and it produces architecturally inconsistent output long before anyone notices.
DIVIDER

4. Engineer the harness

Constrain the agent up front, and verify its work before anyone accepts it. Claude Code hooks are deterministic: they run on every matching event, no matter what the model decides. A PreToolUse hook can block edits to production infrastructure or applied migrations. A PostToolUse hook can format code and scan for secrets after every edit. A Stop hook can run the test suite and hand failures back to Claude to fix before it declares the task done.

Add permission deny rules for secrets and destructive commands, enable the sandbox for OS-level filesystem and network isolation, and define read-only subagents, such as a security reviewer that can inspect a diff but cannot change it. Assign a named owner for this configuration. Harness engineering is an ongoing discipline, and it improves only when someone reviews the telemetry each month.

DIVIDER

5. Redesign oversight by risk tier

Human-in-the-loop for everything does not scale. Human-on-the-loop for everything is a future incident report. Write a delegation matrix instead:

Then make the workflows asynchronous. The Claude Code GitHub Action can respond to a failed CI run by diagnosing the break and opening a fix PR. Claude Code Review can take the first review pass on every PR. Parallel sessions in separate git worktrees let one engineer direct several streams of work without collisions. The developer’s role shifts toward orchestration and judgment, and training should reflect that.

DIVIDER

6. Govern cost and risk in policy

Agents call models and tools across many steps, so controls must sit below the developer. Managed settings let platform teams restrict available models, cap effort levels, disable bypass-permissions mode, allow only managed hooks and MCP servers, and require the sandbox. OpenTelemetry events feed the SIEM for audit. For regulated industries, confirm Zero Data Retention eligibility and cloud deployment through Amazon Bedrock, Google Cloud, or Microsoft Foundry for data residency before the first pilot. Security teams sign off faster when those answers arrive before the questions.

DIVIDER

How UST is putting this to work

UST’s strategic alliance with Anthropic brings Claude into UST’s platforms, engineering services, and internal operations, with a commitment to certify 20,000 UST employees on Claude. The work follows the playbook above. In UST-iDEC, UST’s silicon validation platform, Claude Code serves as the reasoning layer for test scripting and fault detection, where the platform has cut validation cycle times by 50% to 70%.

Within the Alpha AI Practice, we package these patterns for clients: CodeCrafter for legacy modernization, UST-Eval for evaluation harnesses, and ResponsibleRails for red-teaming and governance. The emphasis is the same in every engagement: prove the harness inside our own engineering first, then transfer it with the controls intact.

DIVIDER

A 90-day starting plan

The model will keep improving without your help. The harness will not. That is where the durable advantage sits, and it is the part your organization actually owns.

Ready to move beyond AI coding pilots? Connect with a UST expert to design the governance, harness, and operating model that turns agentic AI into measurable SDLC outcomes.

formId
7e9cb740-6027-49a3-b9de-37c112daede2
portalId
6761677
name
Connect with a UST expert