Français
Applied AI · Advanced 🔴 · Session 3
✏️ Exercises
← Return to program 📄 Source .md

Exercises — Session 3 (Advanced): Claude Agent SDK

Program : Applied AI — Yann Isola Audience : solutions architects — preparation Claude Certified Architect Technical prerequisites: Python ≥ 3.10, pip install claude-agent-sdk ⚠ (check the exact package name in the official documentation), valid API (Application Programming Interface) key.

Each exercise includes: business context, specifications, constraints, evaluation criteria, and commented answer key. Only read the answer key after your attempt.


Exercise 1 — Building an agent with tools (≈ 45 min)

Context

You equip the internal support of a broker with tokenized financial instruments. The agent must answer operators' questions on the status of settlements by querying two internal systems (simulated here by Python functions).

Specifications

  1. Create two tools with the decorator @tool :
    • statut_reglement(reference: str) -> str — returns the status of an operation (simulate with a dictionary: "SETL-001": "réglée", "SETL-002": "en attente de collatéral", "SETL-003": "suspens — incident dépositaire").
    • lister_suspens(gravite: str = "toutes") -> str — lists pending operations, filterable by severity ("critique" / "mineure" / "toutes").
  2. Create a Agent appointed ops-settlement with instructions imposing: response in French, systematic citation of the operation reference, polite refusal of any question outside the payment-delivery scope.
  3. Run via Runner.run() on three test requests:
    • “Where is SETL-002?” »
    • “List me the critical suspenses. »
    • “What is the weather like in Geneva? » (must be refused)

Quality constraints (evaluated)

Evaluation criteria (10 pts)

Criteria Pts
Functional tools, correct diagrams 3
“Prompt engineering” quality docstrings 2
Instructions: scope + effective off-topic refusal 2
Business error returned in text (not raised) 2
The 3 test queries behave as expected 1

Corrected commented

from claude_agent_sdk import Agent, Runner, tool

REGLEMENTS = {
    "SETL-001": {"statut": "réglée", "gravite": None},
    "SETL-002": {"statut": "en attente de collatéral", "gravite": "mineure"},
    "SETL-003": {"statut": "suspens — incident dépositaire", "gravite": "critique"},
}

@tool
def statut_reglement(reference: str) -> str:
    """Retourne le statut d'une opération de règlement-livraison.

    À utiliser dès que l'utilisateur mentionne une référence d'opération
    (format SETL-XXX) et demande son état.

    Args:
        reference: référence de l'opération, ex. "SETL-002".
    """
    op = REGLEMENTS.get(reference.strip().upper())
    if op is None:
        # Erreur MÉTIER → texte actionnable, pas d'exception :
        # le modèle peut rebondir (demander la bonne référence).
        return (f"Aucune opération trouvée pour '{reference}'. "
                f"Vérifier le format (SETL-XXX) ou la date de valeur.")
    return f"{reference} : {op['statut']}"

@tool
def lister_suspens(gravite: str = "toutes") -> str:
    """Liste les opérations en suspens (non réglées).

    À utiliser pour toute demande de vue d'ensemble des suspens,
    éventuellement filtrée par gravité.

    Args:
        gravite: "critique", "mineure" ou "toutes" (défaut).
    """
    if gravite not in ("critique", "mineure", "toutes"):
        return ("Valeur de gravité invalide. "
                "Valeurs acceptées : critique, mineure, toutes.")
    lignes = [
        f"{ref} : {op['statut']} (gravité : {op['gravite']})"
        for ref, op in REGLEMENTS.items()
        if op["gravite"] and (gravite == "toutes" or op["gravite"] == gravite)
    ]
    return "\n".join(lignes) if lignes else "Aucun suspens pour ce filtre."

agent_ops = Agent(
    name="ops-settlement",
    model="claude-sonnet-4-5",  # ⚠ identifiant de modèle volatile
    instructions=(
        "Tu es l'assistant des opérateurs règlement-livraison d'un courtier. "
        "Règles impératives :\n"
        "1. Réponds exclusivement en français.\n"
        "2. Cite TOUJOURS la référence d'opération (SETL-XXX) dans ta réponse.\n"
        "3. Ton périmètre est STRICTEMENT le règlement-livraison. Pour toute "
        "autre demande, décline poliment en une phrase et rappelle ton périmètre."
    ),
    tools=[statut_reglement, lister_suspens],
)

for question in ["Où en est SETL-002 ?",
                 "Liste-moi les suspens critiques.",
                 "Quel temps fait-il à Genève ?"]:
    print(Runner.run(agent_ops, question).final_output, "\n")

Correction points to discuss:


Exercise 2 — Implementing handoffs (≈ 60 min)

Context

One-stop shop for support of a tokenization platform: requests arrive mixed (technical, compliance/KYC — Know Your Customer, billing). You must build a yard officer which routes by handoff to three specialists.

Specifications

  1. Three specialist agents: tech , conformite , facturation — each with instructions clean and at least one dummy tool relevant.
  2. An agent triage :
    • light model (eg. claude-haiku-4-5 ⚠),
    • handoffs towards the three specialists,
    • instructions explicitly prohibiting it from resolving itself.
  3. A callback on_handoff (via the handoff option or a hook) which logs each transfer: triage → conformite (raison).
  4. Tests: three requests, one per specialty, plus one request ambiguous (“My identity verification is blocked and on top of that I was billed twice”) — observe and comment on the choice of model.

Constraints

Evaluation criteria (10 pts)

Criteria Pts
Correct routing of the 3 clear requests 3
Logging of operational handoffs 2
History used (no re-questioning) 2
Return handoff + anti-loop terminal 2
Written analysis of the ambiguous case (5–10 lines) 1

Commented answer key (key extracts)

from claude_agent_sdk import Agent, Runner, handoff, tool

@tool
def verifier_dossier_kyc(client_id: str) -> str:
    """Vérifie l'état du dossier KYC (Know Your Customer) d'un client.

    Args:
        client_id: identifiant client, ex. "C-1024".
    """
    return f"Dossier {client_id} : pièce d'identité expirée, à renouveler."

def log_handoff(source: str, cible: str):
    def _cb(ctx):
        print(f"[HANDOFF] {source} → {cible}")
    return _cb

agent_conformite = Agent(
    name="conformite",
    model="claude-sonnet-4-5",  # ⚠ volatile
    instructions=(
        "Spécialiste conformité/KYC. Tu disposes de l'historique complet : "
        "ne redemande jamais une information déjà fournie. "
        "Si la demande sort de la conformité, retransfère au triage."
    ),
    tools=[verifier_dossier_kyc],
    # handoff retour — ajouté après création du triage (voir note)
)

# ... agents tech et facturation analogues ...

agent_triage = Agent(
    name="triage",
    model="claude-haiku-4-5",   # ⚠ volatile — router = tâche simple, modèle léger
    instructions=(
        "Tu es un aiguilleur. Analyse la demande et transfère au bon "
        "spécialiste : tech (bugs, API), conformite (KYC, identité), "
        "facturation (paiements, factures). Tu ne résous JAMAIS toi-même. "
        "Si plusieurs sujets coexistent, choisis le plus bloquant pour "
        "le client et mentionne l'autre sujet dans ton transfert."
    ),
    handoffs=[
        handoff(agent_tech,        on_handoff=log_handoff("triage", "tech")),
        handoff(agent_conformite,  on_handoff=log_handoff("triage", "conformite")),
        handoff(agent_facturation, on_handoff=log_handoff("triage", "facturation")),
    ],
)

# Handoff retour : câblé après coup pour éviter la référence circulaire.
agent_conformite.handoffs = [handoff(agent_triage,
                                     on_handoff=log_handoff("conformite", "triage"))]

res = Runner.run(agent_triage,
                 "Ma vérification d'identité bloque et j'ai été facturé deux fois.",
                 max_turns=12)   # ⬅ anti-boucle
print(res.final_output)

Correction points:


Exercise 3 — Designing guardrails (≈ 60 min)

Context

Your agent facturation of exercise 2 goes into production in a regulated context. Risk management: “no card data must enter, no promise of financial commitment must go out, and everything must be auditable”.

Specifications

  1. Input guardrail bloquer_pan : detects a PAN (Primary Account Number — bank card number, 16 digits possibly separated by spaces/hyphens) in the incoming message and triggers the tripwire with a message directing to the secure channel.
  2. Output guardrail bloquer_engagement : blocks any output containing a firm promise of reimbursement or a guaranteed amount. Two implementations to deliver and compare:
    • v1: rules/regex (trigger words: “I guarantee you”, “you will be reimbursed for”, amounts + “within X days”…);
    • v2: LLM-as-judge — a lightweight model classifies the output (engagement_ferme / information ), with a judge's prompt that you write.
  3. Audit hooks : class RunHooks journaling on_tool_start/end and guardrails triggers (timestamp, agent name, reason) in a JSONL file (JSON Lines — one JSON object per line).
  4. Application management : catch tripwire exceptions and return a clean fallback message to the user (never stack trace).
  5. Test set : ≥ 6 cases — 2 inputs with PAN (different formats), 1 healthy input, 2 engaging outputs, 1 informative output. Measure false positives/false negatives of v1 vs v2.

Evaluation criteria (10 pts)

Criteria Pts
Input guardrail: detection of the 2 PAN formats, healthy passing input 2
Output guardrail v1 (rules) functional 2
Output guardrail v2 (LLM-as-judge): rigorous judge prompt 2
Encrypted v1/v2 comparison (table FP/FN) + reasoned recommendation 2
JSONL audit hooks + own application fallback 2

Commented answer key (key extracts)

import re, json, datetime
from claude_agent_sdk import (Agent, Runner, RunHooks,
                              input_guardrail, output_guardrail,
                              GuardrailTripwire,
                              InputGuardrailTripwireTriggered,
                              OutputGuardrailTripwireTriggered)

PAN_RE = re.compile(r"\b(?:\d[ -]?){15}\d\b")   # 16 chiffres, séparateurs tolérés

@input_guardrail
def bloquer_pan(ctx, agent, message: str) -> GuardrailTripwire:
    """Refuse tout message contenant un numéro de carte (PAN)."""
    if PAN_RE.search(message):
        return GuardrailTripwire(
            triggered=True,
            message=("Par sécurité, ne transmettez jamais de numéro de carte ici. "
                     "Utilisez le portail sécurisé : Espace client → Paiements."),
        )
    return GuardrailTripwire(triggered=False)

# ---- Output v1 : règles ----
MOTIFS_ENGAGEMENT = [
    r"je vous garantis", r"vous serez rembours",
    r"\d+ ?(€|EUR|CHF).{0,40}sous \d+ jours",
]
@output_guardrail
def bloquer_engagement_v1(ctx, agent, sortie: str) -> GuardrailTripwire:
    """Bloque toute promesse ferme d'engagement financier (règles)."""
    hit = any(re.search(p, sortie, re.IGNORECASE) for p in MOTIFS_ENGAGEMENT)
    return GuardrailTripwire(triggered=hit)

# ---- Output v2 : LLM-as-judge ----
PROMPT_JUGE = """Tu es un contrôleur conformité. Classe le texte suivant :
- "engagement_ferme" : promesse contraignante (remboursement garanti,
  montant précis dû, délai ferme engageant la société).
- "information" : explication de procédure, statut, conditionnel
  ("sous réserve de validation", "généralement").
Réponds par UN seul mot. Texte :
\"\"\"{sortie}\"\"\""""

@output_guardrail
def bloquer_engagement_v2(ctx, agent, sortie: str) -> GuardrailTripwire:
    """Bloque les engagements fermes (juge LLM léger)."""
    verdict = appel_modele_leger(PROMPT_JUGE.format(sortie=sortie))  # ⚠ coût/latence
    return GuardrailTripwire(triggered=verdict.strip() == "engagement_ferme")

# ---- Hooks d'audit JSONL ----
class AuditHooks(RunHooks):
    def _log(self, **champs):
        champs["ts"] = datetime.datetime.utcnow().isoformat()
        with open("audit.jsonl", "a") as f:
            f.write(json.dumps(champs, ensure_ascii=False) + "\n")
    async def on_tool_start(self, ctx, agent, tool):
        self._log(evt="tool_start", agent=agent.name, tool=tool.name)
    async def on_tool_end(self, ctx, agent, tool, result):
        self._log(evt="tool_end", agent=agent.name, tool=tool.name)

# ---- Repli applicatif ----
def repondre(agent, message):
    try:
        return Runner.run(agent, message, hooks=AuditHooks()).final_output
    except InputGuardrailTripwireTriggered as e:
        return e.guardrail_message          # message d'orientation, pas de trace
    except OutputGuardrailTripwireTriggered:
        return ("Votre demande nécessite une validation manuelle. "
                "Un conseiller vous recontacte sous 24 h ouvrées.")

Expected comparison v1 vs v2 (example result):

False positives False negatives Latency Cost
v1 rules Means (blocks “you will be reimbursed if fraud is confirmed) High (paraphrases not listed) ~0ms 0
v2 judge LLM Weak Weak +200–800 ms ⚠ ~1 light call/answer ⚠

Typical recommendation (to be argued): v1 on the front line (fast, free) AND v2 in second line on the outputs that v1 lets pass — defense in depth, judge cost paid only when necessary. Any answer justifying another arbitration (e.g. v2 alone for a low volume flow) is acceptable if the cost/latency/risk reasoning is explicit.


Overall scale