Engineering11 min read

Unrolling the Codex Agent Loop Without Losing Your Mind.

Stop agent loops from rotting: learn the Codex-style driver/state/tool pattern + a working TS loop with streaming tools and guardrails.

Tega Adeyemi
Tega Adeyemi
Unrolling the Codex Agent Loop Without Losing Your Mind.

A practical, code-first guide to the “Codex loop”: streaming tool calls, state, compaction, caching, and production guardrails—so your agents stop vibe-coding and start shipping.

Table of contents

  1. Why “agent loop” matters (and why most implementations rot)
  2. The Codex loop in one sentence (and the 5 moving parts)
  3. Unrolling the loop: a mental model you can implement today
  4. Reference architecture: the Driver, State, Tools, Sandbox, Policy
  5. A minimal loop (TypeScript) that actually works (streaming + tool calls)
  6. Tool design that doesn’t sabotage you (schemas, safety, determinism)
  7. Long-running tasks: compaction + caching (without surprises)
  8. Observability and evals: making the loop measurable
  9. Comparisons: Codex-style loop vs LangGraph vs AutoGen vs CrewAI vs SWE-agent
  10. Production checklist + key takeaways

1) Why “agent loop” matters (and why most implementations rot)

We’ve all seen the demo: an agent edits a file, runs tests, fixes bugs, opens a PR, and everyone claps.

Then reality shows up:

The uncomfortable truth: agents don’t fail because the model “isn’t smart enough.”
They fail because the loop is under-designed.

The OpenAI Codex write-up is valuable precisely because it breaks the magic trick into an implementable system: prompt construction, tool permissions, streaming tool calls, state management, and context control—i.e., the boring parts that decide whether you ship.

2) The Codex loop in one sentence (and the 5 moving parts)

One sentence:

Repeatedly ask the model what to do next, execute the requested tool calls in a constrained environment, feed results back, and stop only when the model produces a final answer.

That sounds obvious—until you implement it and discover there are five separate systems hiding inside:

  1. Driver: orchestrates “ask → tool → observe → continue”
  2. State: what we keep, what we compact, what we discard
  3. Tools: shell, file ops, tests, network, linters, etc.
  4. Sandbox + permissions: what the agent may do (and what it can’t)
  5. Policy & prompt assembly: repo instructions, safety rules, task framing

Codex’s loop emphasizes structured prompt assembly (repo instructions + environment + policy), plus an explicit permissions model for actions/commands. Translation: it’s not “prompt magic,” it’s systems.

3) Unrolling the loop: a mental model you can implement today

Here’s the “unrolled” version we use when building real systems:

Step A — Build the input window

Step B — Ask the model (streaming)

Streaming isn’t just UX candy. It’s operational:

Step C — Execute tool calls (with policy)

Step D — Append observations back into state

Tool output becomes new input items.

Step E — Stop condition

Stop when:

If this feels like building a tiny operating system… yeah. Congrats. You’re now a loop engineer.

4) Reference architecture: Driver, State, Tools, Sandbox, Policy

Driver (the “air traffic controller”)

Responsibilities:

State (the “shared brain”)

Keep:

Decide early:

Important correction: statefulness is a strategy, not a vibe. Pick one:

Mixing both can accidentally double-inject context and inflate costs.

Tools (the “hands”)

Best tools are:

Sandbox + permissions (the “guardrails”)

This is where you prevent:

A sandbox is not optional. It’s your seatbelt. And yes, it’s annoying—like seatbelts.

Policy & prompt assembly (the “adult supervision”)

Our hot take: your best prompt is the one your repo can own. Put it in version control, not in someone’s Notion page that hasn’t been opened since Q2.

5) A minimal loop (TypeScript) that actually works

Below is a corrected “Codex-style” loop using the OpenAI JS SDK + the Responses API.

Key fixes vs the earlier draft:

This is still intentionally minimal. In production you’ll add a real sandbox runner, command allowlists, path policies, timeouts, and structured telemetry.

import OpenAI from "openai";
import { z } from "zod";

const client = new OpenAI({ apiKey: process.env.OPENAI_API_KEY });

// ---------- Tool schema ----------
const RunShellSchema = z.object({
  // In production, don’t accept “any string command”.
  // Prefer structured commands (command + args) and validate hard.
  command: z.string().min(1),
  cwd: z.string().optional(),
});

type ToolResult =
  | { ok: true; stdout: string; stderr?: string; exitCode: number }
  | { ok: false; error: string; exitCode?: number };

// ---------- Sandbox runner (placeholder) ----------
async function runShellSandboxed(args: z.infer<typeof RunShellSchema>): Promise<ToolResult> {
  const { command } = args;

  // ✅ Better than a denylist: allowlist *known safe commands*.
  // This is intentionally tiny; expand per environment.
  const allowed = [
    "npm test",
    "npm run test",
    "npm run lint",
    "pnpm test",
    "pnpm lint",
    "pytest",
    "ruff check .",
    "git diff",
  ];

  const normalized = command.trim();
  if (!allowed.includes(normalized)) {
    return {
      ok: false,
      error: `Command not allowed: "${normalized}". Allowed: ${allowed.join(", ")}`,
      exitCode: 126,
    };
  }

  // 🔥 Replace this with a real sandbox execution (Docker/Firecracker/etc).
  // Return structured output.
  return { ok: true, stdout: `Pretend we ran: ${normalized}\n(ok)`, exitCode: 0 };
}

// ---------- Tool definition (Responses API) ----------
const tools = [
  {
    type: "function" as const,
    name: "run_shell",
    description: "Run an allowlisted command in the project sandbox and return stdout/stderr/exitCode.",
    parameters: {
      type: "object",
      properties: {
        command: { type: "string", description: "Allowlisted command (exact match)." },
        cwd: { type: "string", description: "Optional working directory (restricted in production)." },
      },
      required: ["command"],
      additionalProperties: false,
    },
  },
];

// ---------- Helper types for streaming ----------
type PendingCall = { name: string; argsJson: string };

export async function codexStyleLoop(userTask: string) {
  let previous_response_id: string | undefined;

  // Send initial instructions once, then rely on previous_response_id for statefulness.
  const bootstrapItems: any[] = [
    {
      role: "developer",
      content: [
        {
          type: "input_text",
          text:
            [
              "You are a coding agent operating in a sandbox.",
              "Follow repo instructions if provided.",
              "Be careful: make small changes, run tests when appropriate, and explain decisions briefly.",
              "When you need to run commands, call the run_shell tool with an allowlisted command.",
            ].join("\n"),
        },
      ],
    },
    { role: "user", content: [{ type: "input_text", text: userTask }] },
  ];

  const maxTurns = 12;

  // We keep “new items” per turn. In stateful mode, we don’t resend the entire history.
  let newItems: any[] = bootstrapItems;

  for (let turn = 1; turn <= maxTurns; turn++) {
    const pending: Record<string, PendingCall> = {};
    const toolCalls: Array<{ call_id: string; name: string; arguments: any }> = [];

    let finalText = "";
    let sawCompleted = false;

    const stream = await client.responses.create({
      model: "gpt-5.1-codex-max",
      input: newItems,
      tools,
      stream: true,
      previous_response_id,
    });

    for await (const event of stream) {
      // Stream text output
      if (event.type === "response.output_text.delta") {
        finalText += event.delta;
        process.stdout.write(event.delta);
      }

      // When a function_call output item is created, remember it
      if (event.type === "response.output_item.added" && event.item.type === "function_call") {
        pending[event.item.id] = { name: event.item.name, argsJson: "" };
      }

      // Arguments arrive in deltas
      if (event.type === "response.function_call_arguments.delta") {
        const call = pending[event.item_id];
        if (call) call.argsJson += event.delta;
      }

      // Done = safe point to parse JSON args
      if (event.type === "response.function_call_arguments.done") {
        const call = pending[event.item_id];
        if (call) {
          let args: any;
          try {
            args = JSON.parse(call.argsJson);
          } catch (e) {
            args = { __parse_error: String(e), __raw: call.argsJson };
          }
          toolCalls.push({ call_id: event.item_id, name: call.name, arguments: args });
        }
      }

      if (event.type === "response.completed") {
        previous_response_id = event.response.id;
        sawCompleted = true;
      }
    }

    // If we didn’t get a clean completion, treat it as “incomplete” and decide how to proceed.
    // (In production, this is where you’d apply compaction / retry policies.)
    if (!sawCompleted) {
      return { ok: false, error: "Model response did not complete cleanly (stream ended early)." };
    }

    // Stop if there are no tool calls.
    if (toolCalls.length === 0) {
      return { ok: true, answer: finalText };
    }

    // Execute tool calls and build the next turn’s input
    const nextItems: any[] = [];

    for (const call of toolCalls) {
      if (call.name !== "run_shell") {
        nextItems.push({
          type: "function_call_output",
          call_id: call.call_id,
          output: JSON.stringify({ ok: false, error: `Unknown tool: ${call.name}` }),
        });
        continue;
      }

      const parsed = RunShellSchema.safeParse(call.arguments);
      if (!parsed.success) {
        nextItems.push({
          type: "function_call_output",
          call_id: call.call_id,
          output: JSON.stringify({ ok: false, error: parsed.error.message }),
        });
        continue;
      }

      const result = await runShellSandboxed(parsed.data);
      nextItems.push({
        type: "function_call_output",
        call_id: call.call_id,
        output: JSON.stringify(result),
      });
    }

    // Optional “keep going” nudge (small, safe, avoids runaway)
    nextItems.push({
      role: "user",
      content: [
        {
          type: "input_text",
          text:
            [
              "Continue.",
              "If you changed code, run an allowlisted test/lint command (or explain why not).",
              "If blocked, explain what you need.",
            ].join(" "),
        },
      ],
    });

    newItems = nextItems;
  }

  return { ok: false, error: `Max turns (${maxTurns}) reached.` };
}

What this code demonstrates (for real)

6) Tool design that doesn’t sabotage you

Tool rule 1: Make tools boring

Your model should do the thinking; tools should do the doing.

Good tools:

Bad tools:

Tool rule 2: Prefer diffs over raw edits

Agents are way more reliable when they produce diffs. You can validate:

Tool rule 3: Return machine-parseable output

Even if the model reads it, your driver needs it too:

Tool rule 4: Explicit permissions beat “be careful”

Policy must be executable, not aspirational.

If the agent can run arbitrary shell, it will eventually run:

7) Long-running tasks: compaction + caching

When your agent does real work, the context window grows like sourdough starter. Eventually:

Compaction: shrink the window without losing requirements

OpenAI’s conversation state guidance describes compaction as a way to keep long-running threads manageable when inputs grow. Practically:

Corrected wording: We’re not going to name a specific endpoint here unless you wire it up from the docs—because SDK surfaces and endpoints can evolve. The reliable concept is: compaction exists as a supported pattern, and you should design for it.

Prompt caching: stop paying for the same prefix

Provider-side prompt caching (when available) rewards consistency:

Corrected wording: we’re not claiming a specific “cache bucketing identifier” parameter—just the practical implementation advice: stable prefix good, constantly mutating prefix bad.

8) Observability and evals: making the loop measurable

If you can’t answer these, you don’t have an agent—you have an expensive surprise generator:

Minimum viable telemetry

Log per turn:

Evals that actually matter

For coding agents, use task-level evals:

9) Comparisons: Codex-style loop vs popular frameworks

Let’s keep this spicy and fair.

Codex-style “unrolled loop”

Strengths

Trade-off

LangGraph

LangGraph is designed for graph-based workflows—explicit state, branching, loops.

Great when

Trade-off

AutoGen

AutoGen is a multi-agent conversation framework: roles, chats, humans-in-the-loop, tools.

Great when

Trade-off

CrewAI

CrewAI focuses on orchestrating role-playing agents into a cohesive “crew.”

Great when

Trade-off

SWE-agent

SWE-agent is “agent-loop honest”: it lives in real repos, uses tools, and is evaluation-driven.

Great when

Trade-off

Practical takeaway
Frameworks help you organize the loop. Codex-style unrolling helps you own the loop.

The loop above is one harness among many. Harness engineering: what a harness is (and isn't) gives the general definition, the minimum viable harness for consequential work, and the properties to test any loop against.

10) Production checklist

Loop safety

Reliability

Context control

Observability

Key takeaways

Tega AdeyemiJanuary 26,2026.

The script works. Production is a different sport.

The AI OS letter covers the part tutorials skip: verification, trust, what breaks with real users. One idea, every Saturday, from CM, Cohorte's founder, who has shipped 60+ AI systems.

Free weekly. No spam. Unsubscribe in one click.

Subscribed ✓

The next letter arrives Saturday. Go finish the build.