Ever since the first public version of ChatGPT was released back in 2022, developers have started using it for writing code. While it had some capabilities to produce code, those capabilities were very limited. The tool hallucinated badly and was only good at generating something simple and standard, like boilerplate code and well-known algorithms.

Fast forward to today. Hallucinations are all but gone. There are agentic tools for the full software development lifecycle, so you no longer have to copy snippets from your chat window! And there are teams who no longer manually write any code and don’t end up producing slop!

Of course, agentic development, if not used carefully, can end up producing slop fast, and it often does. However, there are plenty of examples of teams that use it effectively while maintaining high engineering standards. Many of such teams work in environments where quality and precision matter, including the highly regulated industries.

Some examples include Starling Bank, which is one of the biggest digital banks in the UK; Axon, which created AI-enabled mission-critical software and hardware for law enforcement and first responders; and Microsoft, which doesn’t require much introduction. Then there are Airbnb and Uber, which also don’t require much introduction. I can spend hours listing such companies, because there are many.

So yes, you absolutely can still produce clean, performant, and bug-free code by having it all written by agents! If you haven’t experienced it, then you might not believe it. So, stay with me, and I’ll show you. After all, the above-mentioned organizations managed to achieve it!

The trend is clear:

If you want to thrive as a software developer in this new reality, knowing how to use agentic software development tools well is absolutely required. Otherwise, the number of opportunities

But many developers are still catching up to this reality. I mentor developers, and the biggest question I get asked is this:

Where do I start?

This is exactly what we will talk about in this article. I will start by describing the standard steps of the agentic software development process that are applicable to any tool, whether it’s Claude, Codex, Cursor, or something else. They are all pretty much the same these days, so it’s largely a matter of preference.

I will teach you how to use agentic coding tools effectively without turning them into a slop cannon!

Then, I will show you two specific examples of how to apply these steps in Codex: one for a Python app and one for a .NET app.

Finally, I’ll finish the article with a reward for paid subscribers. I will walk you through additional agentic development principles that apply to real production development situations, which include the cleanest way of adding agents to brownfield or legacy codebases, additional caveats to watch for to prevent token spend going haywire, and adding hooks to your agentic flow to automate some steps.

So, let’s begin!

The key steps of the agentic software development process

Let’s start with a hypothetical scenario that is similar to what you may encounter in real life. Imagine you receive a partially specified requirement for a new Python service.

Let’s do a step-by-step walkthrough of how to use the coding agent throughout the lifecycle while retaining explicit engineering control.

1. Start with requirements analysis, not code

Suppose you receive:

Build a Python API that accepts support tickets, classifies them using an LLM, retrieves relevant knowledge-base articles, and proposes a response.

You don’t throw these requirements at the agents and hope for the best. Instead, your first agent instruction could be similar to this:

Analyze these requirements before writing any code. Identify ambiguities, missing requirements, contradictions, assumptions, security concerns, failure scenarios, and decisions that need stakeholder input. Do not implement anything yet.

You want it to discover questions such as authentication, expected load, latency requirements, model/provider choice, data residency, PII handling, response schemas, confidence thresholds, human approval, retries, observability, and what happens when the model or retrieval system fails.

This demonstrates something important: the agent is a requirements reviewer before it is a programmer.

2. Make the agent turn ambiguity into explicit decisions

Here are three important things about how LLMs work that you need to be aware of to use them effectively:

- They have a limited context window

- They work better when the context is not trying to tackle too many things at once

- Old messages in the conversation thread get summarized

The size of the context window is not as much of a problem as it used to be, especially if you use the frontier models. So the model can retain a lot of the conversation you had with it. However, LLMs use their effectiveness when the context contains too many different things.

You can probably remember a situation when you tried to multitask heavily and ended up either doing a very mediocre job at everything you tried or you haven’t even managed to complete anything. LLMs are similar. They work best when their context is targeted.

And the third part is that, even with a large context window, if you have a long-running thread, there will come a point when older messages will become invisible to the LLM. The conversation history will be summarized away to keep within the context window.

And this is why we need to perform the following step:

Don’t simply answer the agent’s questions conversationally and lose the information.

Have it record information in multiple files and the repo, and maintain a structure similar to this:

docs/

requirements.md

decisions.md

architecture.md

implementation-plan.mdThis way, important facts are recorded, so the agent can easily access them later without having to maintain a huge conversation history and keeping everything in its context all the time. Different types of facts are recorded into separate files, so that decisions are separate from requirements, architecture, etc. This helps keep the context targeted.

The content of one such file may look like this:

DEC-004: LLM responses require human approval

Reason: AI-generated customer communication must not be sent automatically.

DEC-005: API p95 target < 3 seconds

DEC-006: PostgreSQL is the system of record

DEC-007: application must remain functional if the LLM provider is unavailableThis creates an audit trail of human decisions rather than leaving important architecture buried in an agent conversation.

3. Ask for architecture options before choosing one

The next step is to brainstorm architecture options with your agent, as you don’t want to end up with an architecture that you will later find to be difficult to work with. You can use a prompt similar to this:

Based on the approved requirements, propose 2–3 implementation approaches. Compare complexity, reliability, testability, operational cost and extensibility. Do not implement anything.

The agent might propose FastAPI + PostgreSQL + direct model SDK versus adding LangChain/LangGraph, for example.

One thing I’d specifically avoid is introducing an agent framework just because the application involves AI.

If the workflow is as simple as this:

request → retrieve → LLM → responseLangGraph may be unnecessary.

If you have something like this:

classify

↓

retrieve

↓

reason

↓

tool call

↓

human approval

↓

retry / branch / resumethen explicit workflow orchestration becomes much more defensible.

That kind of judgment is much more important than knowing lots of agent libraries.

4. Have the agent produce an implementation plan

Once you’ve chosen the architecture, you would prompt your agent to come up with a plan for implementing the features. It would be similar to this:

Enter planning mode. Break the implementation into small, independently testable increments. Specify files/components affected, tests required, acceptance criteria and verification commands for each increment. Do not modify code.

What the agent may come up with may be similar to this (only usually much more detailed):

1. Bootstrap application

2. Define ticket domain model

3. Implement ticket repository

4. Implement classification interface

5. Add deterministic fake classifier

6. Add classification tests

7. Add LLM implementation

8. Add retrieval

9. Add API endpoint

10. Add resilience

11. Add observability

12. Add evaluation suiteThen you review and approve the plan.

This way, you’re preventing an autonomous coding agent from making architectural decisions and immediately propagating them through the repository.

5. Establish the repository’s engineering rules

One of the fastest ways to turn your agentic coding workflow into a slop cannon is to trust agents to make judgment calls on what “good” looks like. This needs to be enforced by a combination of fully deterministic machine-enforced rules and strict agent instructions.

These include things like linters (analyzing your code for excessive complexity, code smells, etc.), static code analysis (analyzing your code for style and structure violations), and other similar tools.

You also need to establish which specific tools you want to use for general development tasks, because the agents still need to execute them to build the application, run the tests, etc.

For Python, the list of suitable tools may include the following:

uv

pytest

ruff

mypy or pyright

pre-commitThen, you will write global agent instructions into a file such as AGENTS.md or CLAUDE.md.

# Engineering rules

## Development

- Python 3.13+

- Use type hints throughout.

- Keep domain logic independent of infrastructure.

- Prefer dependency injection over globals.

- Do not introduce dependencies without justification.

## Testing

- Follow TDD.

- Write a failing test before implementation.

- Implement the minimum code required to pass it.

- Refactor only after tests pass.

- Bug fixes require a regression test.

## AI

- LLM calls must be behind interfaces.

- Unit tests must never require live model calls.

- Structured LLM outputs must be schema validated.

- Treat model output as untrusted input.

- Prompts must be version controlled.

- AI behavior requires evaluation tests.

## Verification

Before declaring work complete:

uv run ruff check .

uv run mypy .

uv run pytestThe distinction I’d emphasize is this:

Agent instructions aren’t controls.

AGENTS.md can tell the agent to run the linter. A deterministic CI pipeline should prevent merging when the linter fails.

That’s much stronger engineering.

6. Use TDD one increment at a time

Whether strict, canonical TDD in legacy software development is the best approach is debatable. However, in agentic development, it absolutely is the best approach. You want your agent to work in small increments (remember context management). And you want your agent to implement the deterministic validation of behavior before it implements the behavior itself.

So, don’t say:

Implement the plan.

Instead, say:

Implement step 3 only using TDD. First add the failing tests and show me the failure. Then implement the minimum production code necessary. Run the relevant tests and quality gates. Stop when this increment satisfies its acceptance criteria.

The loop becomes:

Plan

↓

Human approval

↓

Failing test

↓

Implementation

↓

Tests

↓

Lint/type check

↓

Agent self-review

↓

Human review

↓

Next incrementThis substantially reduces the blast radius of a bad agent decision.

7. Add an adversarial review agent

After one agent implements the plan, you should get another agent, who hasn’t developed any biases based on your continuous conversation history, to review the implementation. So, after implementation, give another agent/review context the requirements and diff and tell it:

Assume this implementation contains subtle defects. Review the diff against the requirements. Look specifically for incorrect assumptions, missing edge cases, security problems, concurrency issues, failure-handling problems, unnecessary complexity and tests that pass without proving the requirement.

You can even separate roles conceptually:

Requirements agent

↓

Planning agent

↓

Human approval

↓

Implementation agent

↓

Test / evaluation system

↓

Review agent

↓

Human approvalThe value comes from separation of generation and verification.

8. For AI applications, go beyond ordinary TDD

AI-enabled software (i.e., software with AI functionality inside it) is different from normal software and is tested differently. Normal software can often be tested as this:

assert calculate_total(...) == 42LLM behavior isn’t usually deterministic enough for that. So the project must have two verification layers.

Deterministic software tests, such as these:

unit tests

integration tests

API contract tests

security tests

failure/retry testsAnd AI evaluations, which verify the following:

golden datasets

classification accuracy

retrieval relevance

structured-output validity

tool-selection correctness

hallucination/factuality checks

prompt-injection tests

latency

token usage

estimated model costFor example, an AI app whose goal is to classify technical support tickets, you might maintain the following list of evaluation datasets:

evals/

classification_cases.jsonl

retrieval_cases.jsonl

adversarial_cases.jsonlEvery significant agent-generated change runs the eval suite. And this includes changing the LLM that the agent inside the app is connected to, even if no agentic code is updated.

You shouldn’t consider an AI-generated implementation correct because another LLM reviewed it. Ultimately, you want executable evidence. That’s what evals provide

9. Make the agent prove completion

Here’s the final part of your local development cycle. At the end of an increment, don’t ask:

Is everything done?

Ask:

Map every acceptance criterion to evidence that demonstrates it is satisfied. Include the relevant test, test result, implementation location and any criterion that cannot currently be proven.

Now the agent is producing evidence rather than confidence statements.

10. Put the same controls into CI

At this point, the local development process is complete. But production development doesn’t stop here. You will now need to add the same check to the continuous integration (CI) pipeline to make sure any violations block the release:

Developer + coding agent

│

▼

Pull request

│

┌─────┴─────┐

▼ ▼

Ruff mypy

│ │

├─────┬─────┤

▼ ▼ ▼

Unit Integration Security

tests tests checks

│

▼

AI evals

│

▼

Agentic review

│

▼

Human review

│

▼

MergeYou give the agent autonomy inside an engineering harness, rather than autonomy over the engineering process.

The agent can investigate, challenge requirements, propose architecture, plan, write tests, implement, refactor, and review. But requirements, architecture decisions, acceptance criteria, and production approval remain explicit human-controlled gates.

So now, let’s look at more specific examples: how do you achieve this by using Codex for Python and .NET apps.

Example of building a Python app in Codex

In the Codex view of the ChatGPT desktop app, I’d implement the workflow almost literally. Codex now has native Plan mode (/plan), project-level AGENTS.md instructions, /review, worktrees, side chats, and the ability to run your repo’s tests/lint/type-check commands.

You can follow these steps:

1. Open the repository in Codex

Use the desktop app → Codex → open your local project folder. Codex is specifically the software-development surface; ChatGPT Work is the more general long-running-work surface.

If starting a new repo, you can use:

/initCodex will use this to scaffold AGENTS.md for the project

2. Give Codex the raw requirements and make it interrogate them

I wouldn’t start coding yet. Paste the interview requirement and say something like this:

Do not implement anything yet.

Act as a senior software architect reviewing these requirements.

1. Read the complete requirement.

2. Identify ambiguities, contradictions and missing information.

3. Identify unstated functional requirements.

4. Identify relevant non-functional requirements.

5. Identify security, privacy, reliability and operational concerns.

6. Identify assumptions that would materially affect the architecture.

7. Separate:

- questions that must be answered before implementation;

- decisions where a reasonable default can be proposed.

Ask me the blocking questions before proceeding.The important interview behaviour here is that you don’t let Codex silently invent requirements.

3. Convert the answers into a specification

After answering its questions:

Update docs/requirements.md with the agreed requirements.

Create docs/decisions.md recording important decisions and their rationale.

Clearly distinguish:

- confirmed requirements;

- assumptions;

- non-goals;

- acceptance criteria.

Do not implement anything.Now your conversation isn’t the source of truth—the repo is.

4. Switch on actual Plan mode

Codex has a native planning mode capability, which can be invoked by typing the following:

/planPlan mode is specifically intended to let Codex investigate the repository and propose an implementation approach before changing code.

Then give it something like this:

Read docs/requirements.md and docs/decisions.md.

Explore the existing codebase.

Develop an implementation plan.

For each milestone specify:

- objective;

- files/components likely to change;

- architectural implications;

- tests required;

- acceptance criteria;

- validation commands;

- dependencies on earlier milestones.

Keep milestones small enough to implement and verify independently.

Do not implement the plan yet.This is much better than merely typing “make me a plan” because you’re using Codex’s actual planning workflow.

5. Challenge its plan

Don’t immediately accept what it produces.

I would say something along these lines:

Before I approve this plan, critique it.

Look for:

- unnecessary complexity;

- premature abstractions;

- unnecessary dependencies;

- missing failure cases;

- security problems;

- concurrency issues;

- testability problems;

- requirements not addressed;

- places where deterministic code would be preferable to an LLM;

- places where an agent/framework is being introduced without justification.

Revise the plan where appropriate.Then you approve it.

6. Create your engineering harness

Before implementation, have Codex establish the Python engineering environment.

For example:

Before implementing application functionality, establish the project's

engineering quality gates.

Use:

- uv for dependency/project management;

- pytest for testing;

- Ruff for linting and formatting;

- mypy for static type checking;

- pre-commit where appropriate.

Configure them using standard Python project conventions.

Add commands for running all quality gates.

Do not implement application functionality yet.

Run the tooling and verify the clean baseline.You might end up with:

pyproject.toml

uv.lock

.pre-commit-config.yamland commands such as:

uv run pytest

uv run ruff check .

uv run ruff format --check .

uv run mypy .7. Put durable rules in AGENTS.md

This is where Codex becomes particularly useful.

AGENTS.md isn’t merely documentation: Codex automatically loads applicable AGENTS.md instructions into its context, including directory-specific instructions.

I’d have something concise like this:

# Engineering instructions

## Development

- Use Python 3.13+.

- Use type hints for production code.

- Prefer simple solutions over speculative abstractions.

- Keep domain logic independent of infrastructure.

- Do not introduce dependencies without justification.

## Development process

Use test-driven development for application behaviour:

1. Write or modify a test demonstrating the required behaviour.

2. Run it and confirm that it fails for the expected reason.

3. Implement the minimum production change required.

4. Run the test.

5. Refactor if necessary.

6. Run the complete relevant test suite.

Bug fixes require regression tests.

## AI

- Keep model-provider code behind interfaces.

- Unit tests must not make real model calls.

- Validate structured LLM output.

- Treat model output as untrusted input.

- Keep prompts version controlled.

- Changes to AI behaviour require relevant evaluation cases.

## Completion

Never claim a task is complete without running the relevant:

- tests;

- Ruff checks;

- type checks.

Report the commands run and their results.There is an important nuance here: OpenAI’s latest Codex guidance recommends avoiding an enormous AGENTS.md; newer models need less micromanagement, so keep durable rules concise and move specialized workflows into skills or task-specific instructions.

8. Implement one planned increment

Now I’d tell Codex:

Implement milestone 1 from the approved plan.

Follow AGENTS.md.

Use TDD.

Do not proceed to milestone 2.

When complete:

1. run the relevant tests;

2. run Ruff;

3. run mypy;

4. inspect the diff;

5. map the milestone acceptance criteria to evidence;

6. report anything that remains unresolved.That “Do not proceed to milestone 2” is useful.

You’re giving the agent considerable freedom within a bounded unit of work.

9. Use Codex’s actual review facility

After the implementation, type the following:

/reviewCodex supports reviewing local changes or comparing them with another branch.

I’d supplement it with:

Review this implementation against docs/requirements.md

and the approved plan.

Prioritize:

- correctness;

- requirement violations;

- security vulnerabilities;

- edge cases;

- concurrency problems;

- error handling;

- tests that do not actually prove the intended behaviour;

- unnecessary complexity.

Do not modify the code during this review.This gives you:

builder → reviewer

rather than:

builder → builder says its own work looks good

10. Use a side chat when you want to challenge something

This is another feature worth showing in an interview.

Codex supports /side, which lets you investigate something without derailing the main implementation thread.

For example:

/side Is PostgreSQL actually justified here, or could this requirement

be satisfied more simply?Or:

/side Explain why you chose LangGraph here. What functionality would

we lose by implementing this as ordinary Python control flow?That’s a very natural senior-engineer workflow.

11. For an AI application, add evals as a first-class quality gate

Tell Codex something like this:

Create an evaluation strategy for the AI behaviour.

Separate deterministic software tests from probabilistic AI evaluations.

Identify appropriate evaluation cases for:

- normal behaviour;

- boundary conditions;

- malformed model responses;

- hallucination;

- retrieval failures;

- incorrect tool selection;

- prompt injection;

- model/provider failure.

Where possible, create a version-controlled evaluation dataset.

Do not change production behaviour until I approve the evaluation design.Then your repository starts looking like this:

src/

tests/

evals/

classification.jsonl

retrieval.jsonl

adversarial.jsonl

docs/

requirements.md

decisions.md

architecture.md

implementation-plan.md

AGENTS.md

pyproject.toml12. Make Codex prove that it’s finished

For the final prompt I’d use something like this:

Perform final verification against docs/requirements.md.

For every acceptance criterion report:

1. requirement;

2. implementation location;

3. test/evaluation proving it;

4. verification result.

Run all relevant:

- unit tests;

- integration tests;

- AI evaluations;

- linting;

- formatting checks;

- static type checks.

Identify any requirement for which there is insufficient evidence.

Do not describe the project as complete if any acceptance criterion

cannot be demonstrated.That’s substantially stronger than his:

Run the tests and make sure everything works.The final workflow becomes similar to this:

Requirements

│

▼

Codex requirements analysis

│

▼

Clarifying questions ◄──── Human answers

│

▼

requirements.md + decisions.md

│

▼

/plan

│

▼

Architecture + implementation plan

│

▼

HUMAN APPROVAL

│

▼

AGENTS.md + quality gates

│

▼

┌──── TDD milestone ─────┐

│ │

▼ │

Failing test │

│ │

Implementation │

│ │

pytest / Ruff / mypy │

│ │

/review │

│ │

Acceptance evidence │

└──────────┬─────────────┘

│

▼

Next milestone

│

▼

Full tests + AI evals

│

▼

Human review

│

▼

PRThat is a true agentically assisted application development rather than just AI-assisted coding.

Building a .NET app agentically

For a .NET application, I’d use essentially the same lifecycle, but make the engineering harness very .NET-native.

1. Give Codex the raw requirement

The original requirement may be similar to this:

Build an ASP.NET Core API that accepts support tickets, uses an LLM to classify them, retrieves relevant knowledge, and generates a suggested response for an agent to approve.

In which case, start with this kind of prompt:

Do not write code yet.

Analyze these requirements as a senior .NET architect.

Identify:

- ambiguities and contradictions;

- missing functional requirements;

- missing non-functional requirements;

- security/privacy concerns;

- reliability and failure scenarios;

- scalability/operational concerns;

- assumptions that would materially affect the design.

Separate:

1. questions that require stakeholder clarification;

2. decisions for which you can propose a reasonable default.

Ask me the blocking questions before proceeding.This part shouldn’t be .NET-specific.

2. Capture the agreed specification

After answering the questions, use a prompt similar to the following:

Create:

docs/requirements.md

docs/decisions.md

Record the agreed requirements, assumptions, non-goals,

acceptance criteria and architectural decisions.

Do not implement anything yet.For an AI application, decisions might include the following:

DEC-001: ASP.NET Core Minimal APIs

DEC-002: PostgreSQL via EF Core

DEC-003: Azure OpenAI behind ITextGenerationService

DEC-004: LLM responses require human approval

DEC-005: LLM outage must not prevent ticket submission

DEC-006: OpenTelemetry for application telemetryNotice that I wouldn’t let Codex immediately decide that Semantic Kernel, Microsoft Agent Framework, LangChain, etc. is required. Make it justify that separately.

3. Enter Plan mode

Use the following command:

/planThen do the following:

Read docs/requirements.md and docs/decisions.md.

Explore the existing solution.

Propose an implementation architecture and incremental plan.

For every milestone specify:

- objective;

- projects/files likely to change;

- interfaces/components introduced;

- tests required;

- acceptance criteria;

- verification commands.

Prefer built-in .NET functionality over additional dependencies.

Do not introduce an agent framework unless the requirements justify it.

Do not implement anything yet.You review that plan before implementation.

4. Have Codex establish the .NET engineering harness

This becomes quite different from Python.

I’d tell Codex the following instructions:

Before implementing application functionality, establish the

engineering quality gates for this .NET solution.

Inspect the repository first and preserve existing conventions.

Configure appropriate:

- compiler/analyzer settings;

- nullable reference types;

- warnings;

- .editorconfig;

- formatting;

- static analysis;

- unit/integration testing;

- code coverage.

Prefer SDK/compiler capabilities before adding third-party tooling.

Do not install global tools unless there is a strong reason.

Keep tooling reproducible through the repository.

Verify the clean baseline before continuing.Codex can inspect what’s already there rather than blindly installing things.

For a modern .NET project, I’d expect some combination of the following:

.editorconfig

Directory.Build.props

Directory.Packages.props

global.json

*.slnx / *.slnand settings such as this:

<PropertyGroup>

<Nullable>enable</Nullable>

<TreatWarningsAsErrors>true</TreatWarningsAsErrors>

<AnalysisLevel>latest</AnalysisLevel>

<EnforceCodeStyleInBuild>true</EnforceCodeStyleInBuild>

</PropertyGroup>Potential verification becomes this:

dotnet restore

dotnet build --no-restore

dotnet test --no-build

dotnet format --verify-no-changesThe exact commands should follow the repo rather than being imposed mechanically.

5. Create AGENTS.md

The /init command will make the agent do it automatically, but it’s up to you to review it and make sure that everything is there. For a .NET project, the structure may look like this:

# Engineering instructions

## .NET

- Follow existing solution conventions.

- Enable and respect nullable reference types.

- Do not suppress compiler/analyzer warnings without justification.

- Prefer framework functionality over unnecessary dependencies.

- Use async APIs for I/O.

- Pass CancellationToken through asynchronous boundaries.

- Do not use .Result or .Wait() on asynchronous operations.

- Keep domain logic independent of infrastructure.

## Architecture

- Keep external services behind abstractions.

- Keep business logic independently testable.

- Do not introduce abstractions solely for hypothetical future requirements.

- Do not introduce an AI/agent framework unless its capabilities are required.

## Testing

Use TDD for application behaviour:

1. Add a failing test.

2. Run it and verify that it fails for the expected reason.

3. Implement the minimum production change.

4. Run the test.

5. Refactor.

6. Run the affected test suite.

Bug fixes require regression tests.

## AI

- LLM providers must be behind interfaces.

- Unit tests must not call real models.

- Validate structured model output.

- Treat model output as untrusted.

- Version-control prompts.

- AI behaviour requires evaluation cases.

## Completion

Before claiming a task is complete:

- build the solution;

- run relevant tests;

- run formatting/static analysis;

- inspect the diff;

- report commands and results.The important thing is not to turn AGENTS.md into a 2,000-line coding standards document.

6. Let Codex implement one vertical slice with TDD

Suppose the first real increment is ticket creation.

Implement milestone 1 only.

Follow AGENTS.md and use TDD.

First write the tests demonstrating the required behaviour.

Run them and confirm that they fail for the expected reason.

Then implement the minimum production code required to make

them pass.

Do not proceed to milestone 2.

When complete:

- run the affected tests;

- build the solution;

- run formatting/static analysis;

- inspect the diff;

- map acceptance criteria to evidence.You might watch Codex produce something like this:

Support.Api

Support.Application

Support.Domain

Support.Infrastructure

Support.TestsBut there’s an important senior-engineering moment here.

If Codex immediately creates the following subfolders:

Domain

Application

Infrastructure

Presentation

Common

Shared

Abstractions

Corefor an application containing three classes, challenge it.

Agentic development doesn’t mean accepting agent-generated architecture.

7. Then implement the AI boundary

For example, rather than sprinkling model calls through controllers, like this:

public interface ITicketClassifier

{

Task<TicketClassification> ClassifyAsync(

SupportTicket ticket,

CancellationToken cancellationToken);

}You can instruct Codex:

Implement the ticket-classification milestone.

The domain/application code must not depend directly on the

chosen model provider.

Use TDD.

Unit tests must use a deterministic fake implementation.

Do not make network/model calls from unit tests.

Structured model responses must be validated before entering

the application domain.Now your architecture might become:

API

│

▼

Application

│

▼

ITicketClassifier

│

├── FakeTicketClassifier ← tests

│

└── AzureOpenAiClassifier ← infrastructure8. Test at the right .NET levels

I’d explicitly ask Codex for several levels of testing:

Unit tests

│

├── domain

└── application

Integration tests

│

├── ASP.NET Core API

├── EF Core

└── PostgreSQL

AI evaluations

│

├── classification

├── retrieval

├── structured output

└── adversarial cases

Architecture/static checksFor ASP.NET Core, WebApplicationFactory<TEntryPoint> is particularly useful for API-level integration tests.

For real infrastructure, Testcontainers can also be justified:

Test process

│

├── ASP.NET Core

│

└── PostgreSQL containerNow you’re testing against PostgreSQL rather than accidentally proving only that EF Core’s InMemory provider works.

9. Have Codex review its implementation

After each meaningful increment, use the built-in review agent:

/reviewThen make the review targeted:

Review the current diff against docs/requirements.md,

docs/decisions.md and AGENTS.md.

Look specifically for:

- incorrect behaviour;

- requirements that are not implemented;

- nullable-reference problems;

- async/await mistakes;

- missing CancellationToken propagation;

- DI lifetime mistakes;

- EF Core query/performance issues;

- concurrency problems;

- security vulnerabilities;

- incorrect HTTP semantics;

- insufficient tests;

- tests that pass without proving the requirement;

- unnecessary abstractions;

- unnecessary dependencies.

Do not modify anything during this review.

Rank findings by severity and provide evidence.That’s much stronger than simply asking, “Does this code look okay?”

10. Add AI-specific evals

This remains important even though the application is .NET. .NET also has agentic frameworks like Microsoft Agent Framework. And these are no worse than Python-based LangChain and LangGraph.

So tell Codex:

Build an evaluation harness for ITicketClassifier.

The evaluation dataset must be version controlled.

Measure at minimum:

- classification accuracy;

- invalid structured responses;

- latency;

- token usage;

- estimated model cost.

Keep these evaluations separate from deterministic unit tests.

Do not make ordinary dotnet test dependent on external model

availability unless the tests are explicitly categorized as

external AI evaluations.That’s an important distinction:

dotnet test

│

└── deterministic

→ PASS / FAIL

AI eval

│

└── probabilistic

→ measurements / thresholdsDon’t contaminate the normal unit suite with unpredictable live LLM calls.

11. Ask Codex to attack the application

I’d add another explicit phase to this:

Act as an adversarial reviewer of the completed AI functionality.

Attempt to identify scenarios involving:

- prompt injection;

- malicious ticket content;

- malformed model responses;

- model timeouts;

- rate limiting;

- provider outage;

- unexpectedly large inputs;

- PII leakage;

- hallucinated knowledge;

- incorrect retrieval;

- cancellation;

- concurrent requests.

For each credible issue, determine whether an automated

regression test can be created.

Do not modify production code until I approve the findings.This is a particularly good use of agentic development because you’re using the model to generate scenarios you may not have considered, not merely code.

12. Final evidence-based verification

I’d finish with this:

Perform final verification against docs/requirements.md.

Create a requirements traceability report.

For every acceptance criterion identify:

1. requirement;

2. implementation location;

3. automated test/evaluation;

4. verification command;

5. result.

Then run all applicable:

dotnet restore

dotnet build

dotnet test

dotnet format --verify-no-changes

Run the AI evaluation suite separately.

Report any requirement that cannot be demonstrated by evidence.

Do not claim that the implementation is complete if any

acceptance criterion remains unverified.The complete .NET Codex flow then becomes this:

RAW REQUIREMENTS

│

▼

Codex gap analysis

│

▼

Clarifying questions

│

HUMAN ANSWERS

│

▼

requirements.md / decisions.md

│

▼

/plan

│

▼

Architecture + plan

│

HUMAN APPROVAL

│

▼

.NET engineering harness

│

┌───────────────────┼──────────────────┐

▼ ▼ ▼

.editorconfig analyzers/build AGENTS.md

│

▼

TDD increment

│

┌────────────┴───────────┐

▼ │

Failing test │

│ │

▼ │

Implementation │

│ │

▼ │

build + test + format │

│ │

▼ │

/review │

│ │

▼ │

Acceptance evidence ──────────────┘

│

▼

AI evaluation suite

│

▼

Adversarial review

│

▼

Requirements traceability

│

▼

HUMAN PR REVIEWThings to watch out for in production

The next section is for paid subscribers only. Inside it, I talk about specific patterns of using software development agents in common production situations.

I start by talking about the process and best practices of integrating agentic tools with brownfield and legacy codebases. Then, I talk about FinOps: the best practices for keeping your token spend low while building real business software. Finally, I describe how to use hooks, which automate certain things in the software development process, so you can save a lot of time and mental effort.

For the rest of you, I’ll wrap up.

I normally write about AI engineering rather than AI-assisted engineering: building AI systems rather than using AI to build systems. This is the first article of the series.

Now, let me know what you think and whether you found this information useful. If you did, then expect me to write more about it in the future.

While it was a long article, we barely scratched the surface. There’s more to agentic AI engineering. Way more.