


Agent Academy
Welcome to Agent Academy, a practical course on building and operating AI agents. It starts from what an agent actually is and ends with running a fleet of them in production, assuming no prior experience.
Work through the eight modules in order. Each one builds on the last, and together they take you from your first agent to a system you can put in front of real users.
Modules
- 1.What an AI agent actually is
- 2.How an agent works, part by part
- 3.Getting a good agent: build, fork or adopt
- 4.Tools, context and capabilities
- 5.Limits, failure modes and honest expectations
- 6.Packaging an agent for the outside world
- 7.Deploying an agent on Yaju
- 8.Operating a fleet with Yaju Agent Systems
Module 1.
What an AI agent actually is
A language model answers. An agent acts. The difference is a loop: the model receives a goal, decides on a single next step, takes that step through a tool, reads what came back, and decides again, repeating until the goal is met, a limit is reached, or it gives up.
That distinction matters more than it sounds. A chatbot produces text and stops. Whatever happens next is done by a person who reads the answer and acts on it. An agent closes that gap itself: it queries the database, opens the pull request, sends the message. The output is not a suggestion, it is a side effect in a real system.
Traditional automation also produces side effects, but it arrives at them differently. A script, an RPA robot or a workflow builder follows a path somebody drew in advance. Every branch is enumerated. Given the same input it does the same thing forever, which is exactly why it is trustworthy, and exactly why it breaks the moment reality steps outside the diagram.
An agent inverts that relationship. You do not write the control flow. You describe the goal, hand over a set of possible actions, and let the model decide the order at runtime. The procedure is generated on the spot, potentially differently each time, shaped by what the agent observes as it goes.
So the real question when choosing between them is not which is more advanced. It is how stable the procedure is. If the steps are known, repeatable and rarely change, write the script: it is cheaper, faster, deterministic and trivially auditable. If the input varies so much that enumerating the branches is impractical, such as messy support tickets, unstructured documents, or investigations where step three depends on what step two found, then an agent starts to earn its cost.
A useful test: take the tools away. If what remains is still useful, you built a chatbot. If it becomes pointless, you built an agent. And if the same behaviour could have been twenty lines of code with a conditional, you probably should have written those twenty lines instead.
Module 2.
How an agent works, part by part
Almost every agent, whatever framework produced it, is assembled from the same handful of parts. Learning to tell them apart is most of debugging, because failures localise cleanly along them.
The model is the reasoning engine. It reads the current state and decides what happens next. Model choice is a real engineering decision rather than a default: a larger model holds up better under ambiguous instructions, a smaller one costs a fraction and replies faster. Mature agents often use more than one, reserving the expensive model for the steps that genuinely need judgement.
The instructions, usually delivered as a system prompt, are a contract rather than a personality. They establish the role, the constraints, the output format and, above all, the definition of done. Vague instructions are the most common single cause of agents that wander in circles or stop halfway.
Context is everything the model can see during this particular step: the task itself, the records you retrieved, the results of earlier tool calls. Treat it as a budget rather than a container. Every token spent on irrelevant material costs money, adds latency, and dilutes the model’s attention on the one line that actually mattered.
Memory is what survives between steps and between runs. Short-term memory is usually just the running transcript. Long-term memory is a deliberate act: a summary you persist, a record of past decisions, a profile of the user. Conflating the two is how agents end up either amnesiac between sessions or slowly drowning in their own history.
Tools are the action space, the typed functions the agent is allowed to call. Their names and descriptions are part of the prompt, because the model selects a tool by reading them. A tool called processData with no description will be chosen badly and blamed on the model.
The decision step is the moment the model emits either a tool call with arguments, or a final answer. Everything upstream exists to make that one choice well.
Execution is the runtime: the component that actually invokes the tool, enforces loop and retry limits, attaches credentials, handles errors and records what happened. Frameworks hand you the reasoning parts. The runtime is what makes the result safe to put in front of other people.
Assembled, the cycle reads: observe the state, choose one action, execute it, observe the result, repeat. When something goes wrong, the symptom usually points at one part. A confidently wrong answer is normally missing context. An inappropriate action is normally a poor tool description. A run that never ends is normally a missing definition of done.
Module 3.
Getting a good agent: build, fork or adopt
There are three honest routes to a working agent, and they trade time against control in predictable ways.
You can write it yourself with an SDK. This gives you complete control over the prompt, the tool surface and the failure handling, and it costs the most time. It is the right choice when the task is specific to your business and no generic version could know your rules. Using a model to help draft the agent is fine and increasingly normal, but treat its output as a first draft from a capable stranger: it does not know your data, your constraints or what a wrong answer would cost you.
You can fork one from a public repository. This is fast, and it is where most people start. It also means inheriting assumptions you did not make, so read before you run. Check what the tools can actually reach, whether any prompt interpolates untrusted input, how credentials are expected to arrive, and whether the licence permits your use. A forked agent that quietly has write access to something important is a bad surprise to discover in production.
Or you can adopt a production-ready agent from a catalogue such as the Yaju Agent Hub, which is built around the open-source rhythm: browse, fork, publish, improve. The distinction from a public repository is that these are agents meant to operate end to end rather than one-off prompt demos, and they arrive already shaped for a governed runtime.
Whichever route you take, the agent is not finished the first time it produces a good answer. Before tuning anything, assemble a small evaluation set: ten to thirty real cases, including the awkward ones you would rather ignore. Write down what a correct response looks like for each, in advance. Judging output after the fact, by impression, is how teams convince themselves an agent improved when it merely changed.
Then iterate deliberately. Change one thing at a time, rerun the whole set, and keep the score. Add every production failure to the set as a regression case, so the same mistake cannot return quietly. Include adversarial cases too: inputs designed to make the agent overstep, invent, or act without enough information.
Finally, personalise it. Your data, your naming conventions, your escalation rules, your definition of done. A generic agent that almost fits is frequently worse than no agent at all, because people abandon it after the second wrong answer and trust is far harder to rebuild than to establish.
And keep retirement on the table. An agent that is not used, or whose task changed underneath it, should be switched off rather than maintained out of sentiment.
Module 4.
Tools, context and capabilities
An agent is only as capable as the actions and information you hand it. This is where most of the practical engineering lives.
Design tools to be narrow and verb-shaped. One tool that reads customer orders beats one that manages the database. Narrow tools produce clear decisions, are easy to permission, and fail in ways you can interpret. Broad tools invite the model to improvise, and the improvisation is what hurts.
Write the description for the model, not for your colleagues. It should say what the tool does, when to use it, when not to use it, and what it returns. Type the arguments and validate them. A schema is not bureaucracy here, it is the difference between a rejected call and a malformed write to a production system.
Resist adding tools speculatively. Selection quality degrades as the action space grows: with thirty tools available the model spends reasoning on choosing rather than on the task, and picks wrong more often. If two tools overlap, merge them or delete one.
Context should be retrieved, not dumped. Fetch the specific records this task needs and pass those. Where the source is large, chunk it sensibly and retrieve by relevance rather than pasting the whole document. Ask for citations back, so an answer can be traced to the passage that produced it.
Separate reading from writing. Most agents need broad read access and very narrow write access, and treating those as one permission is how a summarisation agent ends up able to delete records. Where a write is irreversible, put a human in front of it.
Integrations are where this becomes an organisational problem rather than a coding one. Every team wiring its own connection to the same upstream system produces duplicated credentials, inconsistent policy and no single audit trail. An open standard such as MCP lets any compatible agent reuse the same governed integrations, with one binding per upstream system instead of one per team.
Credentials deserve their own rule, and it is short: the agent never holds a secret. It requests the action, and the platform attaches the credential at the moment of the outbound call. That way a leaked prompt, a verbose log or a shared transcript never becomes a leaked key.
Finally, plan for partial failure. Real systems rate-limit, paginate, time out and return half an answer. An agent that treats every tool response as success will confidently build on nothing. Return errors as structured data the model can reason about, and decide in advance which failures should stop the run rather than be worked around.
Module 5.
Limits, failure modes and honest expectations
Autonomy is the property that oversells itself hardest, so it is worth being precise about what goes wrong.
Errors compound. If each step in a loop is ninety-five per cent reliable, a twenty-step run is only about thirty-six per cent reliable end to end. Worse, a bad assumption at step three does not announce itself: it quietly shapes steps four through twenty, and the final output looks coherent because the agent has been consistent with its own mistake.
Models fabricate. Not randomly, but plausibly, which is the problem. Invented record identifiers, functions that do not exist in an API, confident citations of documents that were never retrieved. Fabrication is most dangerous precisely where verification is most tedious, so design verification in rather than hoping to spot it.
Context failures are subtler than they look. Missing context produces guessing. Stale context produces decisions that were right last quarter. Excessive context produces a model that has technically been told the answer and has not noticed. Each looks like a reasoning failure and none of them is.
Cost behaves non-linearly. A long context is paid for on every single step of the loop, retries multiply it, and a run that fails after fifteen steps costs more than one that succeeds in three. Agents that look cheap in testing routinely surprise people in production, because test cases are short and real ones are not.
Security is genuinely different from ordinary application security. Any content an agent reads can attempt to instruct it, so a support ticket, a web page or a PDF becomes an injection surface. An agent with legitimate credentials that can be talked into misusing them is a confused deputy, and the usual defences do not apply. Constrain what the tools can reach rather than trying to sanitise language.
Maintenance is continuous. Prompts drift out of alignment with a changing business, upstream APIs change shape, models get deprecated or silently updated, and an eval suite that passed in March quietly starts failing in June. An agent is a running system with an owner, not a project that ships once.
All of which leads to the least fashionable conclusion in this course: prefer a few good agents to many mediocre ones. Ten half-working agents cost more, break more often, and erode confidence faster than two that do their job properly. Each additional agent adds surface to monitor, credentials to manage and behaviour to explain. Deciding not to build one, or switching one off, is a legitimate and underused outcome.
Module 6.
Packaging an agent for the outside world
An agent that only runs on the machine where it was written is a prototype. Packaging is what turns it into something another person, team or environment can run without asking you questions.
Start with reproducibility. Pin your dependencies rather than accepting whatever the latest resolves to. Type the inputs and outputs with a schema so callers know exactly what to send and what comes back. Keep configuration outside the code, so the same artefact can run against staging and production without an edit.
Be explicit about what must not travel with it. Secrets, obviously, but also environment-specific endpoints, hard-coded account identifiers and absolute paths. If the agent cannot start without a value, it should fail loudly at startup saying which value is missing, rather than halfway through a run.
Declare the contract the agent needs in order to work: which tools it expects, which permissions those tools require, which model it assumes, and what it costs roughly per run. A reviewer should be able to decide whether to approve it from that declaration alone.
Version it like software, because that is what it is. Keep it in Git, point at a specific version rather than at a branch, diff two versions to see what changed in the prompt as well as the code, and be able to roll back in one command when a change turns out to be worse than it looked.
Ship the evaluation set with the agent. It is the only honest description of what the agent is supposed to do, and it lets whoever inherits it verify that their environment behaves like yours before they trust the output.
Finally, write down the boundaries in plain language: what this agent is for, what it must never do, who owns it, and what to do when it fails. That note costs ten minutes and is the difference between an agent that gets adopted and one that gets quietly abandoned.
Module 7.
Deploying an agent on Yaju
Deployment is the point where an agent stops being yours alone and becomes part of a system other people depend on. The technical act of running it is the easy half.
Production means four things are true that were not true on your laptop. The agent has an identity and a named owner. Something enforces what it is allowed to reach. Everything it did is recorded. And somebody can see what it costs. An agent missing any of those is running in production by accident rather than by design.
On Yaju, the first of those is automatic: an agent gets a unique identity and an owner the moment it runs, so it appears in one catalogue instead of living in somebody’s terminal. This is the practical answer to shadow agents, which is not a policy problem but a visibility one.
From there the Agent Orchestration System attaches what the agent needs at the moment it needs it. Credentials are resolved at runtime from an organisation-level vault, so they never sit in the code or the prompt. Access rules are evaluated before each outbound call, at the level of the workspace and the agent rather than the network. Policy is enforced at the moment of action, not described in a document that nobody reads during an incident.
Observability is on from the first run rather than added after the first surprise. Every model call, tool invocation, input, output and decision is captured as a session trace you can replay, search and export. When an auditor asks what happened on a specific date, the answer is a retrieval rather than an investigation.
For risky actions, put a human in the loop explicitly. The runtime pauses before the irreversible step, requests approval, and continues only once it is granted. This is a design decision that buys a great deal of trust for very little friction, and it is far easier to relax later than to add after an incident.
Deployment itself happens from the command line, which matters more than it sounds: the same commands work inside your existing CI/CD pipelines, so shipping an agent looks like shipping anything else rather than requiring a separate ritual. Where data cannot leave your environment, the same Agent Orchestration System runs self-hosted, in your own cloud or fully air-gapped.
None of this requires changing how the agent was built. Bring your own agent, whatever framework produced it, and the same identity, credentials, policy and tracing apply. That neutrality is deliberate: nobody wants a second layer of lock-in on top of the one they already have with model providers.
In practice, roll out gradually. Run the agent in shadow mode first, where it produces output nobody acts on, and compare against what a person would have done. Then give it a narrow real scope with approval gates. Widen only once the traces are boring.
Module 8.
Operating a fleet with Yaju Agent Systems
One agent is a project. Thirty agents is an operations problem, and it arrives sooner than most teams expect, usually without anyone deciding that it should.
The shift is qualitative rather than quantitative. With one agent you remember what it does. With thirty, spread across teams that each chose their own framework, nobody holds the whole picture, and the questions that matter become unanswerable: what is running, who owns it, what can it touch, what did it do, and what is it costing us.
Inventory comes first, because everything else depends on it. Every agent carries an identity, an owner and a version, whatever produced it. Agents built outside the platform are catalogued on the same terms as those built on it, which is the only way the inventory stays honest.
Control is applied centrally and enforced uniformly. A platform team sets policy once, and the runtime applies it to every team’s agents at the moment of action. This is the arrangement that resolves the usual standoff between engineering, which does not want to ask permission for each agent, and security, which cannot approve what it cannot see.
Observability at fleet scale means more than log retention. Traces let you replay any individual run. An immutable audit trail records the acting user, the credentials used and the outputs produced. Evals score runs continuously against criteria you define, covering behaviour, structure, latency and cost, and a failed critical check halts the agent rather than filing a report about it. That is the step from passive logging to active quality enforcement.
Cost is the dimension teams discover last and regret most. Spend attributed by agent, workspace, user, provider and model, in currency rather than tokens, turns an unexplainable invoice into a list of decisions. Circuit-breaker budgets then make it actionable: warn at eighty per cent, block at a hundred, and stop the overrun instead of reporting it afterwards. Alerts warn you; only a breaker actually stops the spend.
Once costs are visible, they can be reduced systematically rather than anecdotally. Optimisation works across four levers: trying a smaller model, trimming prompts and context, removing tools that add cost without improving outcomes, and moving deterministic steps out of the model and into ordinary code. Each recommendation is measured against your own evals before it is applied, so savings are never bought quietly with quality.
Integrations are governed in the same central way. One binding per upstream system, an approved catalogue teams can request from, and a single audit trail across all of it, rather than each team standing up its own connection and its own copy of the credentials.
Coordination becomes a real concern at scale. Agents hand work to each other, duplicate each other when two teams solve the same problem independently, and compete for the same rate limits. A shared catalogue with reusable agents and skills is the cheapest fix for duplication, and it is also what makes an engineer’s first agent ship in hours rather than weeks.
Finally, treat the fleet as having a lifecycle. Agents get promoted from experiment to production, deprecated when a better one exists, and retired when their task disappears. Measuring adoption per team, cost per agent and performance against evals is what moves the conversation past counting experiments and towards operating on AI rather than merely using it.



