Engineering AI Agents
Bounded tool-using loops with durable state, guardrails, and evals
Learn how an AI agent actually runs: a model choosing actions in a loop that your application code executes, bounds, and stops. This Deep Dive covers least-privilege tool contracts and approval gates, plans that survive failed assumptions, memory with provenance and deletion boundaries, checkpointed execution that recovers without repeating unsafe side effects, and evaluations that turn recorded runs and fault-injected trials into success rates you can trust.
See the Invisible
Interactive simulators visualise what's hidden from view.
Hands-On Labs
Step through executions tick by tick. Manipulate state.
Why, Not Just What
Understand the reasoning behind every design decision.
Quizzes & Cheatsheets
Verify your understanding and keep a quick reference handy.
Get Certified
Earn a shareable certificate to prove your deep expertise.
What's Covered
An agent is a model choosing actions in a loop that application code runs, so the first job is drawing the boundary between model decisions and deterministic application control. You trace one cycle from goal to observation, action, and environment feedback, then bound it with explicit stop reasons and step, token, time, and cost budgets. Planning extends that loop into task decomposition with verifiable subgoals, dependency-aware ordering, and replanning when an observation invalidates an assumption, with progress checks that catch stalled work before the budget runs out.
Tool calls are structured model outputs, and the application decides whether to execute them. That makes the tool portfolio the real safety surface: preconditions and postconditions form the enforceable contract, read-only versus reversible versus irreversible classification decides what needs an approval gate, and least-privilege credentials scoped to each tool cap the damage from a bad call. Prompt injection arriving through untrusted tool observations is contained at that same application-controlled action boundary, and dependencies between parallel actions are checked so conflicting actions never race.
Two kinds of state outlive a single model invocation. Memory splits into working context, episodic records, semantic facts, and procedural memory, each with write criteria, provenance, retrieval tied to the current goal, staleness and contradiction detection, tenant isolation, and retention and deletion boundaries. Execution state lives in an explicit state machine with persisted checkpoints, retry policies by failure class, idempotent actions, duplicate suppression, and compensating actions, so a run that dies halfway resumes without repeating an unsafe side effect or dropping the context a human operator needs to take over.
A single successful run proves little about a nondeterministic system. Evaluation here means asserting task outcomes, constraints, and side effects on recorded runs, then reading trajectories and tool selection alongside step, token, latency, and cost metrics. Stress tests inject tool errors and unavailable dependencies, adversarial cases probe instruction conflicts and unsafe actions, and repeated trial sets give you a measured success rate with a stated confidence, drawn from production traces under privacy controls.
The Curriculum
Comprehensive Lessons! Each with theory, interactive simulation, and quiz.
What an Agent Is: The Execution Loop
Governing Agent Tool Use
Planning Under Uncertainty
Memory That Earns Its Keep
Guardrailed and Recoverable Execution
Evaluating Agent Behavior
Stress Tests and Nondeterministic Trials
This course in one line
The model picks the action, your code decides what runs
Ready to see what's really happening?
All courses included with your subscription. Cancel anytime.