LEARNING PATH · V0.9

A complete AI Engineering path, without forcing you to start at the beginning.

Follow the 10-stage path when you want structure, or enter through a production incident and backfill only the models you actually need.

AI KNOWLEDGE MAP · V1.0

Explore AI as a connected knowledge map.

Open a domain only when you want to inspect the underlying concepts. Every Concept now has a concise explanation; published Guides add a deeper teaching layer. If you want a clear learning order, Courses is the simpler entry point.

01

Understand AI

Build durable mental models for how AI behaves, where it fails, and what software or humans must still guarantee.

12 branches · 44 concepts

AI FoundationsCore mechanics behind model behavior and application constraints.0
Models, Tokens & GenerationHow models turn token sequences into probabilistic outputs.8
Probabilistic model behaviorMENTAL_MODELsharedGuide

Modern language models produce distributions over possible continuations rather than retrieving one fixed answer from memory.

Mental model

Treat every generation as a probabilistic choice conditioned on the current input, context, model parameters, and decoding policy.

Why it matters

The same request can legitimately produce different outputs, so reliable systems need constraints and verification instead of assuming deterministic behavior.

Related concepts

Sampling controls

Read the full Guide
Model claim vs runtime factMENTAL_MODELsharedGuide

A model can suggest what may be true, but only the runtime and external systems can establish what actually happened now.

Mental model

Separate model knowledge and prediction from runtime facts such as tool results, database state, permissions, timestamps, and side effects.

Why it matters

Confusing the two causes hallucinated state, incorrect confirmations, and unsafe actions that the model cannot independently verify.

Read the full Guide
TokenizationCONCEPTsharedGuide

Tokenization converts text or other inputs into discrete units that determine model context length, cost, and sequence structure.

Mental model

Reason about prompts in tokens rather than characters: different languages and strings can consume very different amounts of context.

Why it matters

Token boundaries affect context budgets, truncation, pricing, latency, and sometimes how reliably the model recognizes patterns.

Related concepts

Next-token generation

Read the full Guide
Next-token generationCONCEPTGuide

Next-token generation repeatedly predicts a distribution for the next token, selects one under a decoding policy, appends it, and then predicts again.

Mental model

Generation is an autoregressive loop: each chosen token becomes part of the next input, so early choices can redirect all later probabilities.

Why it matters

This explains why small decoding or context changes can cascade into very different completions and why long outputs accumulate uncertainty.

Related concepts

Tokenization

Read the full Guide
Sampling controlsPRACTICEsharedGuide

Sampling controls such as temperature and probability truncation change how a model selects tokens from its predicted distribution without changing the underlying model weights.

Mental model

Treat decoding settings as output-risk controls: lower randomness narrows choices while higher randomness explores more low-probability alternatives.

Why it matters

Sampling can improve diversity or stability, but it cannot repair missing knowledge, weak reasoning, or incorrect context supplied to the model.

Related concepts

Probabilistic model behavior

Read the full Guide
Context windowCONCEPTsharedGuide

The context window is the finite amount of input and generated history a model can consider within one inference episode.

Mental model

Think of context as working space with a hard capacity: every instruction, example, retrieved passage, tool result, and prior turn competes for room.

Why it matters

Exceeding or poorly allocating the window can silently remove critical evidence and make an otherwise capable model behave inconsistently.

Related concepts

Attention is a budget, not a bag

Read the full Guide
Model capability envelopeMENTAL_MODELsharedGuide

A model capability envelope describes the tasks, modalities, context conditions, reliability ranges, and failure modes within which a model is currently useful.

Mental model

Treat capability as an empirically measured operating region rather than a universal property inferred from a benchmark score or product label.

Why it matters

Systems fail when designers assume a model that succeeds on one workload will generalize to different data, tools, languages, or risk levels.

Related concepts

Capability is not a guarantee

Read the full Guide
Latency and throughput basicsMETRICsharedGuide

Latency measures how long one request takes while throughput measures how much work a system completes over time under a given workload.

Mental model

Analyze first-token latency, total completion time, concurrency, batching, queueing, and resource saturation separately because improving one can worsen another.

Why it matters

AI systems often appear fast in single-user demos but collapse under concurrent traffic when throughput and queueing are ignored.

Read the full Guide
Context & RepresentationHow finite context, ordering and relevance shape model behavior.4
Attention is a budget, not a bagMENTAL_MODELsharedGuide

Attention budget describes the practical competition among tokens and information for a model's limited ability to use context effectively, even before the hard context limit is reached.

Mental model

Treat context as a prioritization problem: high-authority, high-relevance evidence should be easier for the model to use than duplicated, distant, or noisy material.

Why it matters

A prompt can fit inside the context window yet still fail because critical information is diluted among too many competing signals.

Related concepts

Context window · Relevance over raw volume

Read the full Guide
Relevance over raw volumePRACTICEsharedGuide

Context relevance is the degree to which supplied information actually helps the model make the current decision instead of creating noise.

Mental model

More context is not automatically better; each piece should earn its place by changing or supporting the current task outcome.

Why it matters

Irrelevant context consumes attention and budget, increases contradiction risk, and can reduce answer quality even when the window is not full.

Related concepts

Attention is a budget, not a bag

Read the full Guide
Sequence ordering mattersMENTAL_MODELshared

The order of information inside context can change which facts and instructions the model attends to and how it resolves competing signals.

Mental model

Think of context as an ordered sequence with position, adjacency, and recency effects—not as an unordered bag of equally visible facts.

Why it matters

With conflicting instructions or retrieved evidence, changing only the order can change model behavior even when the content is identical.

Context is not memoryMENTAL_MODELsharedGuide

Context is information explicitly supplied for the current inference, while memory is application-managed information selected to persist or be retrieved across interactions.

Mental model

Memory is not a magical extra context window; it is a storage and retrieval system whose selected contents eventually re-enter context when needed.

Why it matters

Separating the two clarifies ownership, persistence, freshness, privacy, and why remembered information can still be absent from a particular model call.

Read the full Guide
Embeddings & MultimodalRepresent meaning across vectors, text, images, audio and other modalities.4
Embeddings as semantic coordinatesCONCEPTsharedGuide

Embeddings map objects such as text, images, or items into vectors whose geometry captures task-relevant patterns of similarity and relationship.

Mental model

An embedding is a learned coordinate system: nearby points are similar according to the training objective, not necessarily equivalent in truth or authority.

Why it matters

Semantic search and clustering depend on this geometry, so retrieval quality is bounded by what the embedding space actually represents well.

Related concepts

Similarity is not relevance

Read the full Guide
Similarity is not relevanceCONCEPTsharedGuide

Vector similarity compares embedding representations to estimate semantic closeness between queries, passages, items, or other encoded objects.

Mental model

Similarity is a ranking signal, not truth: define the embedding space, distance metric, candidate set, and evaluation task before interpreting a score.

Why it matters

High similarity can still retrieve irrelevant or non-authoritative evidence, so downstream filtering and evaluation remain necessary.

Related concepts

Embeddings as semantic coordinates

Read the full Guide
Multimodal representationCONCEPTsharedGuide

Multimodal representation encodes information from text, images, audio, video, or other modalities into forms a model can jointly reason over or transform.

Mental model

Different modalities preserve different evidence; a shared representation enables interaction but can discard spatial, temporal, or structural detail from the source.

Why it matters

Understanding the representation boundary helps decide when a cross-modal summary is sufficient and when the original modality must still be inspected.

Related concepts

Ground claims in the source modality

Read the full Guide
Ground claims in the source modalityMENTAL_MODELshared

Modality grounding ties a claim back to evidence in the original image, audio, video, document, or sensor stream instead of trusting a cross-modal summary.

Mental model

A multimodal representation is an intermediate view; important claims should still be checked against spatial, temporal, or structural evidence in the source modality.

Why it matters

Information can be lost or distorted during modality conversion, especially for precise visual, temporal, or location-sensitive claims.

Related concepts

Multimodal representation

AI Behavior & ReasoningExplore ai behavior & reasoning as part of the AI knowledge map.0
Instructions & PromptingExplore instructions & prompting as part of the AI knowledge map.6
Instruction authority and provenanceMENTAL_MODELsharedGuide

Instruction authority defines which source is allowed to direct behavior when system policy, developer instructions, user requests, retrieved content, and tool data conflict.

Mental model

Assign authority by source role and trust boundary before resolving content; relevance or persuasive wording does not grant permission to override a stronger instruction.

Why it matters

Without a clear hierarchy, untrusted or lower-authority text can redirect the model simply because it is recent or strongly worded.

Related concepts

Ambiguity, specificity and instruction conflict

Read the full Guide
Ambiguity, specificity and instruction conflictRISKsharedGuide

Instruction conflict occurs when multiple instructions cannot all be satisfied simultaneously and the system must resolve them according to authority, scope, and explicit constraints.

Mental model

Detect conflicts explicitly, identify the controlling instruction, preserve non-conflicting requirements, and surface any unresolved constraint instead of silently blending them.

Why it matters

Silent compromise can produce behavior that violates both user intent and system policy while appearing superficially compliant.

Related concepts

Instruction authority and provenance

Read the full Guide
Examples shape behavior; they do not enforce policyMENTAL_MODELshared

Examples can strongly steer format, tone, and local reasoning patterns, but they do not create an enforceable policy boundary.

Mental model

Use examples as behavioral demonstrations; use runtime checks, permissions, and explicit contracts for guarantees that must hold.

Why it matters

This prevents teams from mistaking few-shot imitation for security, authorization, or correctness enforcement.

Prompt vs Context vs Runtime responsibility boundaryMENTAL_MODELsharedGuide

The prompt-context-runtime boundary separates behavioral guidance inside model input from application data and from controls enforced by software outside the model.

Mental model

Use prompts to express intent, context to supply evidence, and runtime code for guarantees such as authorization, state transitions, validation, and side-effect control.

Why it matters

Putting hard guarantees only in prompts converts enforceable system requirements into probabilistic model behavior.

Read the full Guide
Specificity without overconstraintPRACTICEsharedGuide

Prompt specificity makes the requested task, scope, constraints, evidence expectations, and output requirements concrete enough to reduce avoidable ambiguity.

Mental model

Specify the decision boundary rather than adding adjectives: define what to do, what not to do, which evidence matters, and how success will be checked.

Why it matters

Specific prompts improve consistency, but only when they clarify the task; verbosity without sharper boundaries simply consumes context.

Related concepts

Decompose instructions around decisions

Read the full Guide
Decompose instructions around decisionsPRACTICEsharedGuide

Prompt decomposition separates a complex request into smaller reasoning or production steps when intermediate outputs need distinct context, verification, or control.

Mental model

Split only where a boundary changes inputs, evidence, responsibility, or acceptance criteria; do not fragment one coherent task merely to create more prompts.

Why it matters

Good decomposition exposes errors and improves control, while excessive decomposition increases latency, coordination overhead, and context drift.

Related concepts

Planning as search over alternatives · Specificity without overconstraint

Read the full Guide
Reasoning & PlanningExplore reasoning & planning as part of the AI knowledge map.4
Reasoning has a cost and budgetMENTAL_MODELsharedGuide

A reasoning budget limits how much time, tokens, search, tool use, or iterative planning a system spends before deciding, acting, or escalating.

Mental model

Spend additional reasoning only where it can change the decision; define stopping conditions and a simpler fallback for cases where more computation has diminishing returns.

Why it matters

Unlimited reasoning raises latency and cost and can introduce stale assumptions without guaranteeing a better answer.

Related concepts

Verification beats introspection

Read the full Guide
Verification beats introspectionMENTAL_MODELsharedGuide

Verification over introspection means checking outputs against external evidence and executable criteria instead of trusting the model's explanation of its own reasoning.

Mental model

Ask 'what evidence proves this?' before 'why does the model say it believes this?'; prefer tests, source checks, calculations, traces, and independent verifiers.

Why it matters

A model can produce a convincing self-critique while repeating the same hidden mistake, so independent evidence is a stronger correctness signal.

Related concepts

Reasoning has a cost and budget

Read the full Guide
Reasoning opacity and observable evidenceRISKshared

Internal reasoning is not a reliable audit trail; systems should expose observable inputs, actions, outputs, and verification evidence instead.

Mental model

Judge the system through evidence you can inspect and reproduce, not through a persuasive explanation of how it says it reasoned.

Why it matters

Opaque reasoning cannot substitute for traces, tests, source evidence, or runtime facts when debugging or governing a system.

Limits, Uncertainty & FailureExplore limits, uncertainty & failure as part of the AI knowledge map.4
Hallucination as unsupported generationRISKsharedGuide

Hallucination is fluent model output that asserts unsupported, invented, or incorrect information as if it were grounded in available evidence.

Mental model

A language model optimizes plausible continuation, so factual confidence must come from source grounding and verification rather than fluency alone.

Why it matters

The more natural an unsupported statement sounds, the easier it is for users and downstream automation to accept an error without checking it.

Read the full Guide
Uncertainty and calibrationMENTAL_MODELsharedGuide

Uncertainty calibration aligns expressed confidence with observed correctness so confidence levels have operational meaning across many cases.

Mental model

Measure how often predictions are correct at different confidence bands, and route uncertain cases differently rather than treating confidence as decorative prose.

Why it matters

Overconfident errors are especially dangerous in automation because downstream systems may convert a weak guess into a strong action.

Read the full Guide
Distribution shiftRISKsharedGuide

Distribution shift occurs when production inputs, users, environments, or task patterns differ meaningfully from the data and conditions under which a model or evaluation was validated.

Mental model

Treat model quality as conditional on a workload distribution; monitor what changes in the inputs and re-evaluate when the operating population moves.

Why it matters

A model can remain unchanged while real-world accuracy degrades because the environment moved outside the conditions represented in offline testing.

Read the full Guide
Capability is not a guaranteeMENTAL_MODELsharedGuide

Capability describes what a model can often do under favorable conditions, while a guarantee is a property the surrounding system can reliably enforce for every relevant case.

Mental model

Use model capability for flexible cognition and software contracts for invariants; never infer a guarantee merely because a model demonstrated the behavior repeatedly.

Why it matters

Production failures occur when probabilistic competence is mistaken for enforcement of permissions, formats, factuality, or transactional correctness.

Related concepts

Model capability envelope

Read the full Guide
Model LifecycleExplore model lifecycle as part of the AI knowledge map.0
Training & Post-trainingExplore training & post-training as part of the AI knowledge map.4
Pretraining objectiveCONCEPTGuide

The pretraining objective defines the prediction task used to learn broad statistical structure from large datasets before application-specific adaptation.

Mental model

Understand model behavior by asking what training signal rewarded it during pretraining, what patterns that signal teaches well, and what guarantees it never provided.

Why it matters

Many apparent model limitations follow directly from optimizing prediction rather than truth, policy compliance, or task-specific correctness.

Related concepts

Supervised fine-tuning

Read the full Guide
Supervised fine-tuningCONCEPTsharedGuide

Supervised fine-tuning updates model parameters using curated input-output examples so the model more consistently follows a target behavior distribution.

Mental model

Use SFT when the desired behavior can be demonstrated with representative examples and the gap cannot be solved more cheaply at the application layer.

Why it matters

Fine-tuning changes model behavior globally, so poor examples or weak coverage can create persistent regressions that prompting cannot easily isolate.

Related concepts

Preference post-training · Pretraining objective

Read the full Guide
Preference post-trainingCONCEPTsharedGuide

Preference post-training shapes model behavior using comparative or preference signals so outputs better align with desired helpfulness, style, safety, or task behavior.

Mental model

View preference optimization as changing the probability of behaviors, not as installing immutable rules or a factual database inside the model.

Why it matters

Post-training can make a model appear more reliable while still requiring runtime enforcement for permissions, truth, and consequential actions.

Related concepts

Supervised fine-tuning

Read the full Guide
Synthetic data and distillationCONCEPTsharedGuide

Synthetic data and distillation use model-generated examples or teacher outputs to transfer behavior into another training or evaluation dataset.

Mental model

Treat synthetic examples as transformed evidence with inherited biases and errors; validate diversity, correctness, coverage, and leakage before using them as supervision.

Why it matters

Scaling generated data can amplify the teacher's blind spots just as quickly as it expands dataset size.

Read the full Guide
Inference & Model SelectionExplore inference & model selection as part of the AI knowledge map.2
Model selection is a task trade-offPRACTICEsharedGuide

Model selection is a trade-off among capability, latency, cost, context limits, reliability, deployment constraints, and governance requirements.

Mental model

Choose the smallest model that satisfies measured task requirements under the real workload, then escalate only when evidence shows a capability gap.

Why it matters

Defaulting to the largest model can hide architecture problems while increasing cost and latency without materially improving the user outcome.

Related concepts

Open vs closed model trade-offs

Read the full Guide
Open vs closed model trade-offsMENTAL_MODELsharedGuide

Open and closed models trade off control, hosting flexibility, transparency, ecosystem access, capability, operational burden, and vendor dependency differently.

Mental model

Select deployment posture from concrete requirements such as data boundary, customization, latency, cost, auditability, and maintenance capacity rather than ideology.

Why it matters

The wrong choice can create unnecessary infrastructure burden or vendor lock-in without improving the application behavior users actually need.

Related concepts

Model selection is a task trade-off

Read the full Guide
AI Safety & GovernanceExplore ai safety & governance as part of the AI knowledge map.4
Privacy and data boundaryPATTERNsharedGuide

A privacy data boundary defines which personal or sensitive information may enter model context, logs, memory, external tools, or third-party services.

Mental model

Classify data before use, minimize what crosses each boundary, and attach purpose, retention, access, and deletion rules to sensitive information.

Why it matters

AI workflows copy context across many layers, so one poorly controlled transfer can create persistent exposure well beyond the original request.

Related concepts

Human accountability for AI outcomes

Read the full Guide
Bias and fairness as system propertiesRISKsharedGuide

Bias and fairness analysis examines whether an AI system systematically produces different quality, treatment, or risk across relevant groups and contexts.

Mental model

Define the affected groups and decision consequences first, then measure representative slices and investigate the data, model, policy, and workflow causes of disparity.

Why it matters

Aggregate quality can hide concentrated harm, and fairness cannot be established by one universal metric detached from the actual decision context.

Read the full Guide
Human accountability for AI outcomesMENTAL_MODELshared

A person or organization remains responsible for consequential AI outcomes even when generation or decisions are heavily automated.

Mental model

Automation can delegate work, but it cannot erase ownership: define who approves risk, monitors outcomes, and answers for harm or policy violations.

Why it matters

Governance fails when everyone can point to the model while no human role owns the final decision boundary.

Related concepts

Privacy and data boundary

02

Build AI

Engineer AI-native software, knowledge systems, Agents, evaluation and production reliability.

20 branches · 118 concepts

AI-Native SoftwareExplore ai-native software as part of the AI knowledge map.0
Vibe Coding & Agentic CodingExplore vibe coding & agentic coding as part of the AI knowledge map.7
Specification before generationPRACTICEsharedGuide

Specification before generation makes the desired behavior, constraints, acceptance criteria, and non-goals explicit before asking AI to produce an artifact.

Mental model

Write the contract first: define what success means and how it will be checked, then use generation as an implementation step inside that boundary.

Why it matters

Clear specifications reduce prompt thrashing and make reviews distinguish true defects from disagreements about an unstated goal.

Related concepts

AI code review · Plan before code · Tests as executable acceptance evidence

Read the full Guide
Repository context as working memoryPRACTICEsharedGuide

Repository context is the code, architecture, conventions, tests, dependency constraints, ownership, and history an AI coding agent needs to change a real codebase safely.

Mental model

Provide the smallest repository slice that establishes intent and constraints, then let the agent discover additional files through explicit dependency or symbol relationships.

Why it matters

Coding quality depends less on generating syntax than on respecting contracts that exist outside the currently edited file.

Related concepts

Plan before code

Read the full Guide
Plan before codePRACTICEshared

Plan before code resolves intent, affected surfaces, constraints, and verification steps before an AI coding agent edits the repository.

Mental model

A useful plan is a small executable hypothesis about what must change and how you will know the change is safe—not a ceremonial essay.

Why it matters

This reduces broad unnecessary edits and makes architectural regressions easier to catch before code generation compounds them.

Related concepts

Repository context as working memory · Specification before generation

AI code reviewPRACTICEsharedGuide

AI code review evaluates generated or human-written changes against repository intent, architecture, correctness, security, tests, and maintainability rather than judging whether the diff looks plausible.

Mental model

Review from evidence outward: inspect the spec, affected dependencies, executable tests, diff boundaries, and failure cases before accepting the implementation narrative.

Why it matters

AI can produce locally convincing code that violates hidden repository contracts, so review must be independent of the generator's confidence.

Related concepts

Specification before generation · Tests as executable acceptance evidence · Verify dependencies and migrations

Read the full Guide
Tests as executable acceptance evidencePRACTICEsharedGuide

Test-first AI defines executable acceptance evidence before generation or agentic coding so produced changes are evaluated against known behavior instead of aesthetic plausibility.

Mental model

Turn requirements into tests or concrete checks first, then let AI generate toward a fixed target and review any changes to the tests separately.

Why it matters

Without preexisting evidence, an agent can modify both implementation and expectations until its own output appears correct.

Related concepts

AI code review · Specification before generation

Read the full Guide
Verify dependencies and migrationsRISKsharedGuide

Dependency and migration verification checks that library upgrades, schema changes, framework migrations, and generated patches preserve required behavior across affected boundaries.

Mental model

Map what changes transitively, run compatibility and migration tests, inspect generated artifacts, and verify rollback before treating a successful install or build as completion.

Why it matters

Dependency changes often compile while altering runtime semantics, data compatibility, or deployment behavior that local unit tests do not cover.

Related concepts

AI code review

Read the full Guide
Parallel coding agents need isolated work boundariesPATTERNsharedGuide

Parallel agent work runs independent or weakly coupled tasks concurrently when their outputs can be combined without unsafe shared-state conflicts.

Mental model

Parallelize only after identifying true independence, explicit inputs, bounded side effects, and a deterministic join or verification step.

Why it matters

Parallel agents can reduce latency, but hidden dependencies create duplicated effort, conflicting writes, correlated errors, and expensive reconciliation.

Related concepts

Coordination overhead can erase decomposition gains

Read the full Guide
LLM Application EngineeringExplore llm application engineering as part of the AI knowledge map.5
Structured output as an application contractPATTERNsharedGuide

A structured-output contract defines the exact machine-readable shape, types, required fields, and failure behavior expected from model output.

Mental model

Treat model output like an untrusted API response: constrain its schema, parse it, validate semantics, and handle rejection explicitly.

Why it matters

A valid-looking JSON object can still be incomplete or wrong, so schema plus runtime validation is required for dependable automation.

Related concepts

Tool contract and schema design

Read the full Guide
Streaming and backpressureSYSTEM_COMPONENTsharedGuide

Streaming backpressure handles the mismatch between how quickly an upstream AI process produces data and how quickly downstream clients or processors can consume it.

Mental model

Make flow control explicit with bounded buffers, pacing, cancellation, batching, or dropping policies rather than assuming every consumer keeps up indefinitely.

Why it matters

Without backpressure, slow clients can cause memory growth, latency spikes, connection failures, or cascading load across the application.

Related concepts

Queues, concurrency and backpressure

Read the full Guide
Conversation state belongs to the applicationMENTAL_MODELsharedGuide

Conversation state is the application-managed record of turns, task status, user choices, tool outcomes, and other structured information needed to continue an interaction correctly.

Mental model

Keep durable interaction state in explicit software structures and render only the relevant portion back into model context for each turn.

Why it matters

Treating chat history as the state store makes long conversations fragile, expensive, and difficult to validate or resume.

Read the full Guide
Model routing and fallbackPATTERNsharedGuide

Model routing and fallback choose among models or execution paths based on task requirements and recover through predefined alternatives when the preferred path is unavailable or insufficient.

Mental model

Route using observable criteria such as capability need, latency, cost, policy, and failure class; define fallback behavior before an incident occurs.

Why it matters

Ad-hoc switching during failure can silently change quality, context limits, tool behavior, or policy assumptions at the worst possible moment.

Read the full Guide
Rate-limit-aware resiliencePATTERNsharedGuide

API rate-limit resilience keeps an AI application functional when upstream services restrict request volume, concurrency, token usage, or burst rate.

Mental model

Treat rate limits as normal capacity signals: bound concurrency, queue work, back off with jitter, respect retry hints, and degrade or defer noncritical work.

Why it matters

Ignoring limits turns expected capacity pressure into retry storms, long queues, cascading timeouts, and poor user experience.

Read the full Guide
Context EngineeringExplore context engineering as part of the AI knowledge map.5
Finite context budgetMENTAL_MODELsharedGuide

A finite context budget turns prompt construction into a resource-allocation problem across instructions, evidence, memory, and generated history.

Mental model

Allocate context deliberately: reserve capacity for the highest-authority and highest-value information before adding optional detail.

Why it matters

Without an explicit budget, systems tend to accumulate context until important material is truncated, diluted, or too expensive to process.

Related concepts

Context management

Read the full Guide
Context structure and cache-friendly stabilityPATTERNshared

Context can be structured so stable instructions and reusable prefixes stay consistent while volatile task data changes later in the prompt.

Mental model

Separate stable and dynamic context regions; optimize structure first, then let caching exploit the stable prefix where supported.

Why it matters

Cache-friendly structure can reduce cost and latency without sacrificing clarity about context ownership.

Related concepts

Context management

Compaction vs critical-information retentionMENTAL_MODELsharedGuide

Context compaction compresses older or lower-value information into a smaller representation so long-running work can remain within a finite context budget.

Mental model

Compact with a retention objective: decide which facts, decisions, unresolved questions, provenance, and constraints must survive before discarding detail.

Why it matters

Naive summarization can save tokens while deleting exactly the evidence or commitments required for coherent continuation.

Related concepts

Context management

Read the full Guide
Context managementPRACTICEsharedGuide

Context management is the application responsibility for selecting, ordering, compressing, refreshing, and removing information supplied to the model.

Mental model

Manage context as a lifecycle, not a string concatenation step: every item has an owner, freshness, authority, purpose, and removal condition.

Why it matters

Good context management keeps long-running applications coherent while controlling token cost, stale information, and instruction conflicts.

Related concepts

Compaction vs critical-information retention · Context structure and cache-friendly stability · Finite context budget · Retrieval pipeline

Read the full Guide
Prompt caching and stable prefixesPATTERNshared

Prompt caching reuses computation for stable prompt prefixes; it is a performance optimization, not a substitute for context management.

Mental model

Design stable prefixes around genuinely reusable instructions and data, then measure hit rate, invalidation behavior, cost, and latency.

Why it matters

Poor caching can couple prompts to provider details or preserve stale context while delivering little real savings.

Knowledge SystemsExplore knowledge systems as part of the AI knowledge map.0
RAG & RetrievalExplore rag & retrieval as part of the AI knowledge map.7
Evidence granularity and chunkingPRACTICEsharedGuide

Evidence granularity determines how large or small retrieved units are and therefore how precisely a claim can be supported without losing necessary surrounding context.

Mental model

Choose chunks around the decision: small enough to isolate relevant evidence, large enough to preserve definitions, relationships, and provenance needed to interpret it.

Why it matters

Chunks that are too large add noise, while chunks that are too small separate facts from qualifiers and create misleading retrieval.

Related concepts

Retrieval pipeline

Read the full Guide
Dense, sparse and hybrid retrievalSYSTEM_COMPONENTsharedGuide

Hybrid retrieval combines complementary retrieval signals such as lexical matching and vector similarity instead of trusting one ranking mechanism alone.

Mental model

Use each retriever for the failure mode it handles well, then fuse or rerank candidates under one evaluation objective.

Why it matters

Exact identifiers and semantic paraphrases behave differently, so combining signals can improve recall without giving up precision on literal evidence.

Related concepts

Retrieval pipeline

Read the full Guide
Reranking as a second evidence-selection stageSYSTEM_COMPONENTGuide

Reranking applies a stronger scoring step to a smaller retrieved candidate set so the most useful evidence is ordered more accurately for the final context.

Mental model

Use cheap retrieval for broad recall, then spend more computation on a bounded candidate set where better relevance judgment can change the final evidence order.

Why it matters

Reranking can improve precision without making first-stage retrieval expensive, but it cannot recover documents that were never retrieved at all.

Related concepts

Retrieval pipeline

Read the full Guide
Freshness, authority and source priorityMENTAL_MODELsharedGuide

Source authority defines which evidence is allowed to establish a fact when multiple documents, memories, or generated statements disagree.

Mental model

Assign authority before retrieval: distinguish primary sources, governed internal records, user-provided context, secondary commentary, and model-generated claims.

Why it matters

Without authority rules, a system can retrieve relevant but weaker evidence and confidently overwrite the source that should control the decision.

Related concepts

Retrieval pipeline · Source-backed drafting

Read the full Guide
Structured and metadata-aware retrievalPATTERNsharedGuide

Metadata retrieval uses structured attributes such as source, date, tenant, document type, permissions, product, or jurisdiction to constrain and rank evidence before generation.

Mental model

Use metadata for facts the system already knows exactly, then combine it with semantic retrieval for meaning that cannot be captured reliably as fields.

Why it matters

Semantic similarity alone can return the wrong customer, version, time period, or access scope even when the text looks highly relevant.

Read the full Guide
Retrieval pipelineSYSTEM_COMPONENTsharedGuide

A retrieval pipeline turns a query into candidate evidence through indexing, query transformation, retrieval, filtering, reranking, and context assembly stages.

Mental model

Make each retrieval stage observable and evaluable so recall, precision, authority, latency, and cost failures can be attributed to the right component.

Why it matters

Treating RAG as one opaque vector search call hides where evidence was lost and makes quality improvements difficult to reproduce.

Related concepts

Context management · Dense, sparse and hybrid retrieval · Evaluate retrieval separately from generation · Evidence granularity and chunking · Freshness, authority and source priority · Reranking as a second evidence-selection stage

Read the full Guide
Evaluate retrieval separately from generationPRACTICEsharedGuide

Retrieval evaluation measures whether a retrieval system finds the right evidence for representative queries before generation quality is considered.

Mental model

Evaluate retrieval separately with labeled or inspectable evidence sets, slice by query type, and distinguish recall failures from ranking or authority failures.

Why it matters

A strong generator cannot answer from evidence that retrieval never supplied, so end-to-end scores alone can hide the real bottleneck.

Related concepts

Retrieval pipeline

Read the full Guide
Memory SystemsExplore memory systems as part of the AI knowledge map.5
Working memory vs durable memoryMENTAL_MODELsharedGuide

Working memory supports the current task, while durable memory persists selected information across sessions or workflows under explicit governance.

Mental model

Do not persist everything: promote information from working to durable memory only when it has clear value, scope, authority, and retention rules.

Why it matters

Separating the two prevents temporary guesses and transient task state from becoming long-lived facts that contaminate future behavior.

Read the full Guide
Memory write/read lifecycle and selection policySYSTEM_COMPONENTsharedGuide

Memory lifecycle governs when information is captured, written, scoped, retrieved, corrected, refreshed, expired, and deleted across an AI system.

Mental model

Every memory entry needs provenance, owner, scope, write policy, freshness expectations, and an explicit end-of-life condition.

Why it matters

Persistent memory becomes a long-lived source of behavior, so stale or incorrectly scoped entries can repeatedly corrupt future decisions.

Related concepts

Memory vs authoritative source-of-truth conflict

Read the full Guide
Memory vs authoritative source-of-truth conflictRISKsharedGuide

Memory authority conflict occurs when stored user or system memory disagrees with newer, more authoritative, or differently scoped evidence.

Mental model

Resolve memory by provenance, scope, freshness, and source authority; remembered content should never silently override a stronger current source.

Why it matters

Persistent memory can make an outdated preference or earlier mistake repeatedly win over corrected facts unless conflicts are handled explicitly.

Related concepts

Expiry, invalidation, versioning and forgetting · Memory write/read lifecycle and selection policy

Read the full Guide
Expiry, invalidation, versioning and forgettingPATTERNsharedGuide

Memory expiry removes or invalidates stored information when its retention period, freshness window, authority, consent, or task relevance no longer holds.

Mental model

Every durable memory should answer 'when does this stop being safe to use?' through TTL, source revision, explicit deletion, or policy-driven invalidation.

Why it matters

Memory without expiry turns obsolete preferences and facts into persistent hidden context that can silently bias future outputs.

Related concepts

Memory vs authoritative source-of-truth conflict

Read the full Guide
User memory vs domain knowledge boundaryMENTAL_MODELshared

User memory stores person-specific state and preferences, while domain knowledge represents shared facts that should come from governed sources.

Mental model

Keep personalization and source-of-truth knowledge in separate authority domains with different write, expiry, and correction policies.

Why it matters

Mixing them can let one user's remembered preference overwrite authoritative knowledge or leak into another user's context.

Search & Knowledge GraphsExplore search & knowledge graphs as part of the AI knowledge map.2
Knowledge graph structure and relation-aware lookupSYSTEM_COMPONENTsharedGuide

Knowledge graph structure represents explicit entities and relationships so systems can reason over known connections that are difficult to recover reliably from text similarity alone.

Mental model

Use graph structure when identity and relationships are first-class facts; use unstructured evidence for rich descriptions, then connect both without duplicating source authority.

Why it matters

Explicit relations make multi-hop, dependency, lineage, and constraint queries more dependable than hoping embeddings infer every structural connection.

Related concepts

Combine structured and unstructured search

Read the full Guide
Agent EngineeringExplore agent engineering as part of the AI knowledge map.0
Tools & Function CallingExplore tools & function calling as part of the AI knowledge map.5
Tool contract and schema designPATTERNsharedGuide

A tool contract defines the operation name, parameters, types, side effects, permissions, error surface, and result semantics exposed to an AI agent.

Mental model

Design tools as narrow APIs for uncertain callers: make valid actions easy, invalid actions impossible or explicit, and consequential effects visible before execution.

Why it matters

Ambiguous tool interfaces force the model to guess hidden application semantics, increasing invalid calls and unsafe side effects.

Related concepts

Capability boundary and least privilege · MCP protocol, application state, authorization and capability boundaries · Structured output as an application contract · Tool-result validation and explicit error surfaces

Read the full Guide
Tool-result validation and explicit error surfacesPRACTICEsharedGuide

Tool-result validation checks that external tool responses are structurally valid, semantically plausible, authorized, and suitable for the next action.

Mental model

A successful HTTP or function call is only transport success; validate the returned facts, identifiers, status, and error conditions before trusting them.

Why it matters

Agents often fail after a tool call appears successful but returns stale, partial, malformed, or unexpected data that the model then treats as authoritative.

Related concepts

Tool contract and schema design

Read the full Guide
Capability boundary and least privilegePATTERNsharedGuide

Least privilege gives an AI agent only the minimum capabilities, data access, scope, and duration required to complete its current responsibility.

Mental model

Grant narrow permissions to specific operations and resources, then expand only when a demonstrated task requirement cannot be satisfied safely otherwise.

Why it matters

Model mistakes and prompt injection become far less damaging when the runtime simply does not expose unnecessary powers.

Related concepts

Runtime enforcement vs model persuasion · Tool contract and schema design

Read the full Guide
Reversible vs irreversible actionsMENTAL_MODELsharedGuide

Reversible actions can be safely undone or corrected, while irreversible actions create side effects whose consequences cannot be reliably rolled back.

Mental model

Classify action reversibility before granting autonomy; use previews, drafts, staging, soft deletes, and approval gates to postpone irreversible commitment.

Why it matters

The acceptable level of model uncertainty is much lower when an action moves money, deletes data, sends messages, or changes external state permanently.

Related concepts

Human approval at the high-risk boundary

Read the full Guide
Human approval at the high-risk boundaryPATTERNsharedGuide

A human review boundary defines which AI-generated decisions or actions require explicit human approval before they can create consequential effects.

Mental model

Place review at transitions where uncertainty and impact become unacceptable, and give the reviewer the evidence needed to approve, reject, or revise the action.

Why it matters

Human-in-the-loop is ineffective when approval is ceremonial, arrives too late, or lacks the context required to catch the actual risk.

Related concepts

Reversible vs irreversible actions

Read the full Guide
MCP & Capability InterfacesExplore mcp & capability interfaces as part of the AI knowledge map.3
MCP protocol, application state, authorization and capability boundariesMENTAL_MODELsharedGuide

MCP boundaries separate protocol messages and capability descriptions from application state, authorization, user consent, and runtime enforcement that the protocol alone does not own.

Mental model

Use MCP to describe and transport capabilities, while keeping permission, business state, validation, and consequential action policy in the host application.

Why it matters

A protocol connection does not grant trust or authority; confusing transport with permission exposes tools more broadly than intended.

Related concepts

MCP discovery, routing and capability interfaces · Tool contract and schema design

Read the full Guide
MCP discovery, routing and capability interfacesSYSTEM_COMPONENTGuide

MCP capability negotiation lets participants discover supported features and interfaces so they can choose compatible interactions without assuming every server or client implements the same surface.

Mental model

Discover first, then route only to capabilities explicitly advertised and understood by both sides; treat missing capability as a normal compatibility state.

Why it matters

Protocol evolution creates heterogeneous deployments, so hard-coded assumptions about available features lead to brittle integrations.

Related concepts

MCP protocol, application state, authorization and capability boundaries · MCP Tasks extension and multi-round-trip interaction boundaries

Read the full Guide
MCP Tasks extension and multi-round-trip interaction boundariesSYSTEM_COMPONENTshared

MCP Tasks and elicitation-style interactions extend a single request/response tool call into multi-round-trip work with explicit lifecycle and user-input boundaries.

Mental model

Treat long-running task state, elicited information, cancellation, and completion as protocol/application responsibilities rather than hidden conversational state.

Why it matters

This version-sensitive boundary matters when MCP workflows become asynchronous or require structured user participation across multiple turns.

Related concepts

MCP discovery, routing and capability interfaces

Agent RuntimeExplore agent runtime as part of the AI knowledge map.8
Act → observe → verify loopSYSTEM_COMPONENTGuide

An agent loop repeatedly observes state, decides what to do, acts through available capabilities, and evaluates whether to continue or stop.

Mental model

Model the agent as an explicit state machine with observation, decision, action, verification, and stopping conditions rather than an open-ended chat.

Why it matters

The loop is where retries, permissions, memory, side effects, and recovery interact, so leaving it implicit makes failures difficult to bound.

Related concepts

Close the loop with verification · Loop vs Graph responsibility boundary · Workflow decomposition

Read the full Guide
Planning vs direct executionMENTAL_MODEL

Planning and direct execution are two runtime strategies: planning adds an intermediate decision process, while direct execution acts from the current state immediately.

Mental model

Choose planning only when decomposing or comparing alternatives can materially improve the next action; otherwise prefer the simpler direct loop.

Why it matters

Unnecessary planning increases latency, token cost, and stale assumptions without necessarily improving outcomes.

Related concepts

Bounded autonomy, termination and escalation

Bounded autonomy, termination and escalationPATTERNsharedGuide

Bounded autonomy lets an agent act independently only inside predefined capabilities, budgets, risk limits, and escalation conditions.

Mental model

Autonomy is a permission envelope, not a personality trait: define what the system may do, how far it may continue, and when a human must take over.

Why it matters

Useful agents need freedom to act, but unbounded freedom converts model uncertainty into real operational or financial consequences.

Related concepts

Planning vs direct execution

Read the full Guide
Checkpoint, resumable state and recoveryPATTERNsharedGuide

Checkpoint and resume persist enough explicit workflow state that a long-running AI task can recover after interruption without replaying unsafe or completed work.

Mental model

Checkpoint durable state at meaningful boundaries: completed steps, outputs, external side effects, pending decisions, version metadata, and the next valid transition.

Why it matters

Long-running agents inevitably encounter crashes and timeouts, so recovery must continue from known state rather than reconstructing intent from conversation text.

Related concepts

Async and long-running task lifecycle

Read the full Guide
Interrupt and cancellation semanticsPATTERNGuide

Interrupt and cancellation semantics define how a long-running AI task is asked to stop, what work may still complete, and how partial state or side effects are reconciled.

Mental model

Cancellation is a state transition, not a promise that all work instantly disappears; record the request, propagate it, observe acknowledgements, and reconcile already-started effects.

Why it matters

Without explicit semantics, users can believe a task stopped while external operations continue or retries later revive abandoned work.

Related concepts

Async and long-running task lifecycle

Read the full Guide
Async and long-running task lifecycleSYSTEM_COMPONENTsharedGuide

Asynchronous long-running work separates task initiation from completion so operations can outlive a single request, connection, or model turn.

Mental model

Represent long work with durable task identity, lifecycle state, progress, cancellation, result retrieval, retries, and ownership outside the chat request.

Why it matters

Without explicit async semantics, users and agents cannot safely distinguish still-running work from failed work or decide whether a retry is appropriate.

Related concepts

Checkpoint, resumable state and recovery · Interrupt and cancellation semantics

Read the full Guide
Workflow decompositionPRACTICEsharedGuide

Workflow decomposition splits a complex objective into explicit stages with clear inputs, outputs, dependencies, and verification boundaries.

Mental model

Decompose around decisions and evidence, not around arbitrary prompt count; each stage should have a reason to exist and a checkable handoff.

Why it matters

Good decomposition localizes failures, enables parallelism where safe, and prevents one oversized model call from hiding multiple responsibilities.

Related concepts

Act → observe → verify loop · Long-form consistency

Read the full Guide
Close the loop with verificationPATTERNsharedGuide

A verification loop repeatedly compares a produced result with explicit acceptance evidence and uses failures to drive bounded revision or escalation.

Mental model

Generate → verify → classify failure → revise or stop; verification must have independent criteria and a finite budget rather than asking the same model to 'try again'.

Why it matters

Iteration only improves quality when failure information changes the next action; unstructured retries can repeat the same mistake indefinitely.

Related concepts

Act → observe → verify loop

Read the full Guide
Computer & Browser UseExplore computer & browser use as part of the AI knowledge map.3
Observe the environment before actingPRACTICEshared

Environment observation reads the current visible or machine-readable state before an agent decides what action is valid.

Mental model

Use an observe → interpret → act loop; never let a stale plan stand in for checking the state that the action will actually affect.

Why it matters

Interfaces change and actions have side effects, so stale state can make an otherwise correct browser or computer-use plan unsafe.

Related concepts

Ground actions in visible state

Ground actions in visible statePATTERNsharedGuide

Action grounding binds an agent's intended operation to the actual current environment state, target object, and available affordances before execution.

Mental model

Resolve 'what exactly will this action affect?' from fresh observation and stable identifiers, then validate that the intended target still matches the visible state.

Why it matters

A correct plan can become unsafe when interfaces change, elements move, or state becomes stale between planning and execution.

Related concepts

Observe the environment before acting

Read the full Guide
Sandbox and permission boundariesPATTERNsharedGuide

Sandbox permissions limit which files, networks, processes, credentials, tools, and external side effects an AI coding or computer-use agent can access during execution.

Mental model

Design the sandbox as the real capability boundary: default deny, grant task-specific access, isolate credentials, and require stronger approval when crossing into higher-risk resources.

Why it matters

Prompt instructions cannot reliably contain a compromised or mistaken agent if the runtime already exposes unrestricted host capabilities.

Read the full Guide
Multi-Agent & OrchestrationExplore multi-agent & orchestration as part of the AI knowledge map.5
Loop vs Graph responsibility boundaryMENTAL_MODELsharedGuide

Loop versus graph is the architectural choice between one agent repeatedly deciding the next action and an explicit workflow whose branches, joins, and stages are predetermined by application structure.

Mental model

Use a loop for flexible local decisions and a graph when dependencies, parallelism, approvals, or recovery boundaries deserve explicit topology.

Why it matters

Turning every task into a graph creates overhead, while forcing structured workflows into one open loop hides control and makes recovery harder.

Related concepts

Act → observe → verify loop · Sequential, conditional and parallel topology trade-offs

Read the full Guide
Sequential, conditional and parallel topology trade-offsMENTAL_MODELGuide

Orchestration topology determines how loops, agents, stages, branches, joins, and shared state are connected, creating different latency and failure-propagation trade-offs.

Mental model

Choose the simplest topology that matches real dependencies, then evaluate coordination cost, isolation, retry scope, and verification at joins.

Why it matters

A more elaborate multi-agent graph can look sophisticated while producing slower, more correlated, and harder-to-debug failures than a single controlled loop.

Related concepts

Delegation contract and shared vs isolated state · Loop vs Graph responsibility boundary

Read the full Guide
Delegation contract and shared vs isolated statePATTERNGuide

Delegation state records what work was assigned to another agent, with which inputs, constraints, authority, status, outputs, and ownership for the eventual result.

Mental model

Delegation is a stateful contract, not a message: preserve task identity, expected result, dependencies, progress, and who verifies or integrates the output.

Why it matters

Without explicit state, multi-agent systems lose track of duplicate work, stale assignments, partial results, and responsibility at joins.

Related concepts

Independent verification, coordination cost and correlated failure · Sequential, conditional and parallel topology trade-offs

Read the full Guide
Independent verification, coordination cost and correlated failureMENTAL_MODELshared

Independent verification separates the producer of an answer or action from the mechanism that checks it, reducing shared assumptions and correlated failure.

Mental model

A verifier should receive the evidence and acceptance criteria it needs without inheriting every hidden assumption of the producer.

Why it matters

Adding more agents only improves reliability when their failures are not perfectly correlated.

Related concepts

Coordination overhead can erase decomposition gains · Delegation contract and shared vs isolated state

Coordination overhead can erase decomposition gainsRISKGuide

Coordination overhead is the extra latency, communication, state management, verification, and failure handling introduced when work is split across multiple agents or stages.

Mental model

Treat coordination as a cost that must be earned by better decomposition, parallelism, specialization, or independent verification compared with a simpler baseline.

Why it matters

Multi-agent systems often add more messages and moving parts than useful intelligence, making them slower and less reliable without clear benefit.

Related concepts

Independent verification, coordination cost and correlated failure · Parallel coding agents need isolated work boundaries

Read the full Guide
Production AIExplore production ai as part of the AI knowledge map.0
Evaluation & ReliabilityExplore evaluation & reliability as part of the AI knowledge map.10
Timeout ambiguity: missing response is not confirmed failureRISKsharedGuide

A timeout means the caller did not receive a response in time; it does not prove that the remote operation failed or never produced a side effect.

Mental model

Treat timeout as an unknown outcome until status, idempotency state, or an authoritative system of record resolves what actually happened.

Why it matters

Blindly retrying an ambiguous timeout can duplicate payments, refunds, messages, or other irreversible actions.

Related concepts

Idempotency boundary for repeated intent · Retry policy and retry amplification

Read the full Guide
Retry policy and retry amplificationRISKsharedGuide

Retry amplification occurs when retries multiply load, duplicate side effects, or trigger more failures instead of simply recovering from a transient error.

Mental model

Retry only errors that are safe and likely transient, with bounded attempts, backoff, idempotency, and visibility into the original operation's outcome.

Why it matters

During incidents, uncontrolled retries can turn a small dependency failure into a system-wide outage or repeated financial action.

Related concepts

Compensation and recovery after side effects · Timeout ambiguity: missing response is not confirmed failure

Read the full Guide
Idempotency boundary for repeated intentPATTERNsharedGuide

An idempotency boundary ensures that repeated delivery of the same logical intent does not create repeated external side effects.

Mental model

Assign a stable intent identity before the irreversible operation and make every retry converge on the recorded outcome of that identity.

Why it matters

Retries are unavoidable in distributed systems, so idempotency is the mechanism that keeps transport uncertainty from becoming duplicated business actions.

Related concepts

Timeout ambiguity: missing response is not confirmed failure

Read the full Guide
Compensation and recovery after side effectsPATTERNsharedGuide

Compensation and recovery handle partially completed workflows by repairing, reversing where possible, or explicitly reconciling side effects that cannot simply be retried away.

Mental model

After a failure, first determine what actually committed, then choose resume, compensate, reconcile, or escalate based on the side effect's reversibility and authority.

Why it matters

Distributed workflows rarely fail atomically, so assuming 'error means nothing happened' can duplicate or corrupt real-world state.

Related concepts

Retry policy and retry amplification

Read the full Guide
Traceability as causal execution historyMENTAL_MODELsharedGuide

Traceability links a system outcome back through the sequence of inputs, model calls, retrievals, tool actions, state changes, and decisions that produced it.

Mental model

Preserve causal execution history with stable identifiers so an operator can move backward from an observed failure to the responsible step and evidence.

Why it matters

Without traceability, debugging becomes guesswork and evaluation cannot tell whether a failure came from the model, retrieval, tool, or runtime.

Related concepts

Observability as a diagnosis interface

Read the full Guide
Evaluation environment and verifier designMENTAL_MODELsharedGuide

Evaluation evidence is the observable data used to judge whether an AI system satisfies a specific quality, safety, or release requirement.

Mental model

Start from the decision you need to make, then collect test cases, measurements, traces, and verifier outputs that can actually support or block that decision.

Why it matters

Without explicit evidence, teams mistake anecdotes, model confidence, or a few successful demos for proof that the system is ready.

Related concepts

Cost, latency, quality and release vetoes as one decision · Dataset slices, regression and aggregate-improvement traps · Outcome vs trajectory evaluation

Read the full Guide
Outcome vs trajectory evaluationMENTAL_MODELshared

Outcome evaluation judges the final result; trajectory evaluation also inspects the actions, tool calls, and decisions used to reach it.

Mental model

Use outcome metrics for end-state quality and trajectory evidence when unsafe, wasteful, or brittle paths can still produce a superficially correct result.

Why it matters

Two agent runs can end at the same answer while one violates policy, wastes resources, or relies on an unrecoverable path.

Related concepts

Evaluation environment and verifier design

Dataset slices, regression and aggregate-improvement trapsMENTAL_MODELGuide

Dataset slices reveal whether an apparent aggregate improvement hides regressions for specific user groups, languages, tasks, risk categories, or operating conditions.

Mental model

Define meaningful slices before release and require important slices to satisfy their own thresholds rather than accepting only a higher overall average.

Why it matters

AI changes often redistribute quality, so a global metric can improve while a critical cohort becomes significantly worse.

Related concepts

Confidence, variance and sample-size humility · Evaluation environment and verifier design

Read the full Guide
Confidence, variance and sample-size humilityMENTAL_MODELsharedGuide

Confidence and variance describe how uncertain an observed metric or model evaluation is across samples, judges, runs, or dataset slices.

Mental model

Read every score together with sample size, dispersion, confidence interval, and slice behavior rather than treating one average as exact truth.

Why it matters

Small or noisy evaluations can reverse apparent winners, causing teams to ship changes whose measured improvement is mostly random variation.

Related concepts

Dataset slices, regression and aggregate-improvement traps

Read the full Guide
Cost, latency, quality and release vetoes as one decisionMETRICsharedGuide

Release economics evaluates quality, latency, inference cost, operational risk, and business value together when deciding whether an AI change is worth shipping.

Mental model

Translate technical metrics into one decision surface with thresholds and vetoes; an improvement is useful only if its total cost and risk fit the product objective.

Why it matters

A model can score better while making the product slower, more expensive, or operationally fragile enough that the release is still a bad trade.

Related concepts

Evaluation environment and verifier design

Read the full Guide
ObservabilityExplore observability as part of the AI knowledge map.3
Observability as a diagnosis interfaceMENTAL_MODELsharedGuide

Observability is useful when telemetry lets an operator explain why a system behaved as it did, not merely when many metrics are collected.

Mental model

Design logs, traces, metrics, and model/tool evidence around diagnostic questions so failures can be localized across system layers.

Why it matters

Without diagnosis-oriented observability, teams see that quality dropped or latency rose but cannot identify the causal component to fix.

Related concepts

Attribute failures across model, retrieval, tool and runtime layers · Traceability as causal execution history

Read the full Guide
Attribute failures across model, retrieval, tool and runtime layersPRACTICEsharedGuide

Failure attribution identifies whether an observed AI-system failure originated in the model, context, retrieval, memory, tool, data, policy, runtime, or interaction between layers.

Mental model

Trace the failure through observable boundaries and isolate the first layer whose evidence diverges from expected behavior before changing the whole system.

Why it matters

Fixing the wrong layer creates prompt patches for data problems, model upgrades for tool bugs, and expensive changes that leave the original cause intact.

Related concepts

Observability as a diagnosis interface

Read the full Guide
Online monitoring and drift signalsSYSTEM_COMPONENTsharedGuide

Online monitoring watches production behavior after release using operational, quality, safety, cost, and drift signals tied to real traffic.

Mental model

Define what would indicate degradation before shipping, instrument those signals, and connect alerts to investigation, rollback, or traffic-control actions.

Why it matters

Offline evaluation cannot represent every production distribution shift, so release is the beginning of evidence collection rather than the end.

Related concepts

Bounded rollout, fallback and graceful degradation

Read the full Guide
Security & Human ControlExplore security & human control as part of the AI knowledge map.4
Trust boundary and defense in depth for untrusted contextMENTAL_MODELsharedGuide

A trust boundary marks where data, instructions, identities, or capabilities cross from one authority domain into another and must be validated before gaining influence.

Mental model

Assume anything crossing the boundary is untrusted until provenance, authorization, structure, and allowed effect are checked by the receiving layer.

Why it matters

Prompt injection and confused-deputy failures occur when untrusted content is allowed to inherit privileges it never legitimately possessed.

Related concepts

Prompt injection requires runtime defense

Read the full Guide
Prompt injection requires runtime defensePATTERNsharedGuide

Prompt injection defense prevents untrusted content from gaining instruction authority or reaching capabilities beyond the data role it was meant to play.

Mental model

Separate data from instructions, preserve provenance, restrict tools by policy, validate requested actions, and enforce consequential boundaries outside the model.

Why it matters

Prompt wording alone cannot reliably stop malicious retrieved text or user content from attempting to redirect an agent with real capabilities.

Related concepts

Trust boundary and defense in depth for untrusted context

Read the full Guide
Runtime enforcement vs model persuasionMENTAL_MODELsharedGuide

Runtime enforcement implements non-negotiable constraints in software that can validate, allow, deny, transform, or require approval before an AI-driven action proceeds.

Mental model

Move guarantees out of persuasive text and into executable gates around tools, data, permissions, state transitions, and side effects.

Why it matters

Models can misunderstand or ignore instructions probabilistically, but production systems still need deterministic boundaries around consequential behavior.

Related concepts

Capability boundary and least privilege

Read the full Guide
Auditability for consequential actionsPATTERNsharedGuide

Auditability is the ability to reconstruct what inputs, policies, model versions, tools, approvals, and actions produced a consequential outcome.

Mental model

Record the minimum evidence needed for an independent reviewer to trace the decision without relying on hidden memory or a participant's recollection.

Why it matters

When incidents, disputes, or governance reviews occur, systems that cannot reconstruct their decisions cannot reliably explain or improve them.

Read the full Guide
Production ArchitectureExplore production architecture as part of the AI knowledge map.5
Cross-layer architecture decompositionMENTAL_MODEL

Cross-layer architecture separates model behavior, context/retrieval, tools, runtime state, policy, and product guarantees so failures can be assigned to the layer that can control them.

Mental model

Do not solve every problem in the prompt; map each requirement to the layer that has the authority and observability to enforce it.

Why it matters

Production AI failures often come from responsibility leaking between layers rather than from one isolated model error.

Related concepts

Bounded rollout, fallback and graceful degradation

Bounded rollout, fallback and graceful degradationPATTERNGuide

Rollout and fallback control how a new AI behavior receives production traffic and how the system returns to a known safer state when evidence turns negative.

Mental model

Use staged exposure, explicit success and veto signals, and a tested rollback or fallback path that does not depend on the failing component itself.

Why it matters

A safe release process limits blast radius and gives teams time to learn before a model or architecture change reaches every user.

Related concepts

Cross-layer architecture decomposition · Evidence synthesis into SHIP / BLOCK / INCONCLUSIVE · Online monitoring and drift signals

Read the full Guide
Evidence synthesis into SHIP / BLOCK / INCONCLUSIVEMENTAL_MODELsharedGuide

SHIP, BLOCK, and INCONCLUSIVE distinguish evidence that supports release, evidence that violates a veto, and evidence that is insufficient to make a defensible decision.

Mental model

Do not force every evaluation into pass/fail; define release thresholds and vetoes, and preserve an explicit state for uncertainty that requires more evidence.

Why it matters

Treating missing evidence as success encourages risky releases, while treating it as automatic failure can block learning when the honest answer is simply unknown.

Related concepts

Bounded rollout, fallback and graceful degradation

Read the full Guide
Queues, concurrency and backpressureSYSTEM_COMPONENTshared

Queue backpressure controls how producers, workers, and downstream services behave when incoming work exceeds safe processing capacity.

Mental model

Model queue depth, concurrency limits, admission control, retry behavior, and cancellation together so overload slows or sheds work instead of cascading.

Why it matters

Uncontrolled concurrency turns latency spikes into retry storms, duplicate work, and resource exhaustion in long-running or agent workloads.

Related concepts

Streaming and backpressure

Version external capabilities explicitlyPATTERNsharedGuide

Versioned dependencies make external libraries, models, protocols, APIs, and data schemas explicit inputs whose changes can alter system behavior.

Mental model

Pin what must be reproducible, record compatibility assumptions, and test migrations as behavior changes rather than treating upgrades as routine housekeeping.

Why it matters

AI systems depend on fast-moving components, so silent upgrades can change prompts, tokenization, tool contracts, latency, or evaluation outcomes.

Read the full Guide
Model EngineeringExplore model engineering as part of the AI knowledge map.3
Fine-tuning dataset designPRACTICEsharedGuide

Fine-tuning dataset design chooses representative examples, labels, balance, difficulty, negative cases, and evaluation separation needed to teach a target behavior without importing avoidable bias or leakage.

Mental model

Design the dataset around behavior gaps and decision boundaries, then preserve held-out evidence that the model never trains on for honest evaluation.

Why it matters

Fine-tuning quality is bounded by dataset quality; more examples can reinforce the wrong behavior if coverage, labels, or leakage are poorly controlled.

Read the full Guide
Adapter fine-tuning and parameter-efficient adaptationSYSTEM_COMPONENTsharedGuide

Adapter fine-tuning updates a small set of additional or selected parameters, such as LoRA adapters, to specialize a base model without retraining all model weights.

Mental model

Use adapters as one deployment option in a broader adaptation decision: measure task gains, serving complexity, data quality, compatibility, and rollback needs.

Why it matters

Parameter efficiency lowers training cost but does not remove dataset risk, evaluation requirements, or operational complexity at inference time.

Read the full Guide
Inference serving, quantization and deployment trade-offsSYSTEM_COMPONENTshared

Inference serving turns a chosen model into an operational service by managing memory, batching, quantization, concurrency, latency, and deployment constraints.

Mental model

Treat serving as a systems problem: the same model can have very different cost and latency profiles depending on runtime, hardware, batching, and precision choices.

Why it matters

Model quality alone does not determine whether an adapted model is practical to deploy or economical to operate.

03

Use AI

Apply AI to real outcomes such as writing, research, knowledge work and business workflows.

13 branches · 45 concepts

Create with AIExplore create with ai as part of the AI knowledge map.0
Long-Form CreationExplore long-form creation as part of the AI knowledge map.5
Long-form consistencyRISKGuide

Long-form consistency keeps facts, terminology, voice, references, and structural commitments coherent across many sections generated over time.

Mental model

Maintain explicit canonical notes and review checkpoints outside the model context so later chapters can be checked against stable decisions.

Why it matters

Long documents drift gradually, and small contradictions across chapters can undermine trust even when each paragraph looks individually strong.

Related concepts

Workflow decomposition

Read the full Guide
Research-to-outline workflowPRACTICEsharedGuide

Research-to-outline converts collected evidence into a structured argument or chapter plan where each section has a purpose, supporting sources, unresolved questions, and logical relationship to the whole.

Mental model

Build the outline from claims and evidence rather than from generic headings; every section should answer why it exists and what source-backed point it must establish.

Why it matters

Moving directly from research notes to drafting encourages repetition, missing evidence, and chapters whose structure follows generation convenience instead of reasoning.

Related concepts

Decompose a research question

Read the full Guide
Source-backed draftingPRACTICEshared

Source-backed drafting ties factual claims to known sources while the text is written instead of adding citations after unsupported prose already exists.

Mental model

Draft from an evidence set and preserve claim-to-source links so later revision can update or remove claims when sources change.

Why it matters

Fluent unsupported text becomes expensive to verify once it has propagated across chapters and arguments.

Related concepts

Freshness, authority and source priority

Editorial voice and style controlPRACTICEsharedGuide

Editorial voice control keeps generated writing aligned with explicit tone, terminology, audience, style, and consistency rules across many outputs.

Mental model

Represent voice as reviewable constraints and examples, then check produced text against those rules instead of relying on vague instructions like 'sound professional'.

Why it matters

Stable voice is essential for long-form and product content because small stylistic drift accumulates across sections and weakens trust.

Read the full Guide
Fact-check and revision loopPRACTICEsharedGuide

Fact-check and revision separates drafting from evidence verification, then uses identified support gaps or contradictions to make targeted corrections.

Mental model

Extract checkable claims, trace each to authoritative sources, classify support, and revise only after the evidence status is explicit.

Why it matters

Asking a model to 'review its draft' often preserves the same unsupported assumptions; claim-level evidence breaks that coupling.

Read the full Guide
Media CreationExplore media creation as part of the AI knowledge map.3
Visual briefing before generationPRACTICEGuide

A visual brief translates a creative goal into explicit subject, composition, hierarchy, style, constraints, references, and acceptance criteria before image generation.

Mental model

Specify what the image must communicate and how it should be judged before describing decorative details or choosing a generation model.

Why it matters

Without a brief, iteration becomes subjective prompt tweaking and teams cannot distinguish a model failure from an unclear visual objective.

Read the full Guide
Multimodal iteration loopPRACTICEsharedGuide

Multimodal iteration improves an artifact by alternating between generation, direct inspection of visual or audio evidence, targeted changes, and explicit quality checks.

Mental model

Inspect the actual modality after each meaningful change; textual descriptions of the asset are not substitutes for seeing or hearing the produced result.

Why it matters

Many visual and audio defects are obvious in the artifact but invisible in prompts or metadata, so inspection must remain part of the loop.

Read the full Guide
Media quality reviewPRACTICEsharedGuide

Media quality review evaluates generated visual, audio, or multimedia assets against communication intent, technical correctness, consistency, accessibility, and publication constraints.

Mental model

Review the actual rendered asset at target size and channel using explicit criteria; metadata and prompts describe intent but cannot prove the final media succeeded.

Why it matters

Generation pipelines can produce technically valid files that still contain visual errors, unreadable text, inconsistent branding, or inaccessible presentation.

Read the full Guide
Courses & Knowledge ProductsExplore courses & knowledge products as part of the AI knowledge map.3
Learning objective designPRACTICEGuide

Learning objective design states what a learner should be able to understand, decide, produce, or verify after a learning experience and under what conditions.

Mental model

Write objectives as observable capability changes, then align content, practice, and assessment to the same decision or behavior.

Why it matters

Without clear objectives, AI-generated courses can look comprehensive while failing to produce any measurable change in learner behavior.

Related concepts

Curriculum decomposition

Read the full Guide
Curriculum decompositionPRACTICEGuide

Curriculum decomposition organizes a learning goal into a sequence of prerequisite concepts, decisions, practice, and transfer tasks that build usable capability progressively.

Mental model

Decompose by learner dependency rather than content category: each unit should prepare a concrete ability required by a later task or decision.

Why it matters

A large content library is not a curriculum if learners encounter advanced ideas before the mental models and practice needed to use them.

Related concepts

Learning objective design · Package and publish a knowledge product

Read the full Guide
Package and publish a knowledge productPRACTICEshared

Knowledge-product publishing packages reviewed learning content into a versioned, navigable, distributable artifact with clear ownership and release criteria.

Mental model

Publishing is a release step: validate completeness, media, rights, navigation, versioning, and update ownership before distribution.

Why it matters

A good curriculum can still fail as a product when learners cannot navigate it, trust its version, or receive maintained updates.

Related concepts

Curriculum decomposition

Knowledge WorkExplore knowledge work as part of the AI knowledge map.0
AI Knowledge BaseExplore ai knowledge base as part of the AI knowledge map.2
Knowledge-base lifecycleSYSTEM_COMPONENTsharedGuide

A knowledge-base lifecycle governs how knowledge is ingested, normalized, indexed, refreshed, corrected, expired, evaluated, and eventually removed.

Mental model

Treat a knowledge base as a maintained product with ownership and freshness policies, not as a one-time upload of documents into a vector store.

Why it matters

Even excellent retrieval degrades when the underlying knowledge becomes stale, duplicated, contradictory, or ownerless.

Related concepts

Customer-support copilot and escalation

Read the full Guide
Curate, update and retire knowledgePRACTICEshared

Knowledge curation continually decides what enters a knowledge base, what stays authoritative, what must be updated, and what should be retired.

Mental model

Treat ingestion as an editorial lifecycle with ownership, provenance, freshness, duplication, and deletion rules—not as a one-time upload job.

Why it matters

Retrieval quality cannot remain high when the corpus accumulates stale, conflicting, or ownerless material.

Research with AIExplore research with ai as part of the AI knowledge map.4
Decompose a research questionPRACTICEGuide

Research question decomposition turns a broad question into answerable subquestions with explicit definitions, evidence needs, scope, and unresolved assumptions.

Mental model

Decompose by claims that would change the final conclusion, then map each subquestion to the evidence required to support or falsify it.

Why it matters

Without decomposition, AI research tends to collect broadly relevant information without proving the specific claims needed for a defensible conclusion.

Related concepts

Claim–evidence matrix · Research-to-outline workflow

Read the full Guide
Triangulate independent sourcesPRACTICEsharedGuide

Source triangulation compares independent evidence streams before accepting a claim, especially when any single source may be incomplete or biased.

Mental model

Seek sources with different failure modes, then record where they agree, conflict, or leave the claim unresolved instead of averaging them blindly.

Why it matters

Multiple citations add little value when they repeat the same upstream error; independence is what makes corroboration meaningful.

Related concepts

Claim–evidence matrix

Read the full Guide
Claim–evidence matrixPATTERNsharedGuide

A claim-evidence matrix maps each important claim to the sources that support, contradict, qualify, or fail to resolve it, making research reasoning auditable.

Mental model

Treat claims and sources as a many-to-many structure: record evidence strength, independence, freshness, and disagreement instead of attaching one convenient citation per paragraph.

Why it matters

The matrix exposes unsupported conclusions and correlated sources before they become polished prose that is difficult to challenge.

Related concepts

Decompose a research question · Synthesize without erasing disagreement · Triangulate independent sources

Read the full Guide
Synthesize without erasing disagreementPRACTICEshared

Research synthesis combines evidence into a conclusion while preserving uncertainty, source disagreement, and unresolved questions instead of averaging them away.

Mental model

Separate supported consensus, conflicting evidence, and open gaps; synthesize the evidence structure before writing a single narrative conclusion.

Why it matters

A smooth summary can hide the exact disagreements that should change a decision or trigger further research.

Related concepts

Claim–evidence matrix

Documents & Data AnalysisExplore documents & data analysis as part of the AI knowledge map.3
Extract structure from documentsPRACTICEsharedGuide

Document extraction converts files such as PDFs, images, tables, or forms into structured evidence while preserving where each value came from.

Mental model

Treat extraction as a data pipeline with source coordinates, parsing confidence, schema validation, and recoverable errors rather than as free-form summarization.

Why it matters

Downstream analysis cannot be verified if extracted numbers, fields, or citations lose their connection to the original document.

Related concepts

Structure analysis before synthesis

Read the full Guide
Structure analysis before synthesisPRACTICEshared

Structured analysis separates extraction, normalization, calculation, interpretation, and synthesis so each transformation can be checked independently.

Mental model

Turn messy material into explicit tables, fields, assumptions, and intermediate results before asking the model for a final narrative.

Why it matters

Unverifiable reasoning often enters when raw evidence and interpretation are collapsed into one opaque generation step.

Related concepts

Extract structure from documents · Verify AI-assisted data analysis

Verify AI-assisted data analysisPRACTICEsharedGuide

Data-analysis verification independently checks extraction, transformations, calculations, assumptions, and uncertainty before accepting an AI-assisted analytical conclusion.

Mental model

Separate the pipeline into source data → structured data → calculation → interpretation, and verify each transition with reproducible evidence.

Why it matters

A fluent analytical narrative can hide a single extraction or arithmetic error that invalidates the final business conclusion.

Related concepts

Structure analysis before synthesis

Read the full Guide
Work AutomationExplore work automation as part of the AI knowledge map.2
Choose what to automate and what to keep humanMENTAL_MODELsharedGuide

The workflow automation boundary decides which steps should run automatically and which decisions still require explicit human or system approval.

Mental model

Automate repeatable, observable, reversible work first; keep high-uncertainty or high-consequence transitions behind stronger review gates.

Why it matters

Automation creates leverage only when the cost of mistakes is bounded; otherwise it simply accelerates the propagation of bad decisions.

Related concepts

Automation needs ownership and maintenance

Read the full Guide
Automation needs ownership and maintenanceMENTAL_MODELsharedGuide

Automation maintenance is the ongoing work required to keep an AI workflow correct as APIs, prompts, models, data, business rules, and user behavior change.

Mental model

Price automation by lifecycle cost, not setup effort: include monitoring, exception handling, dependency updates, evaluation, ownership, and manual recovery paths.

Why it matters

An automation that saves minutes today can become operational debt if nobody owns its failures or understands how to repair it later.

Related concepts

Choose what to automate and what to keep human

Read the full Guide
Business with AIExplore business with ai as part of the AI knowledge map.0
Product, Marketing & GrowthExplore product, marketing & growth as part of the AI knowledge map.2
Customer-problem research with AIPRACTICEsharedGuide

Customer problem research gathers evidence about recurring user situations, pains, alternatives, willingness to change, and existing behavior before building or automating a solution.

Mental model

Research the problem independently of your proposed product: look for repeated costly behavior, current workarounds, decision triggers, and evidence that users already care.

Why it matters

AI makes building cheap, which increases the risk of efficiently producing solutions for problems that customers do not value enough to adopt or pay for.

Related concepts

Content marketing as a repeatable system

Read the full Guide
Content marketing as a repeatable systemSYSTEM_COMPONENTshared

A content-marketing system connects customer questions, reusable content production, distribution, measurement, and feedback into a repeatable operating loop.

Mental model

Treat content as an owned pipeline with inputs, cadence, channels, conversion signals, and maintenance—not as isolated AI-generated posts.

Why it matters

For a solo business, repeatability and feedback compound while one-off generation quickly becomes an unsustainable queue of assets.

Related concepts

Customer-problem research with AI

Support & OperationsExplore support & operations as part of the AI knowledge map.1
Customer-support copilot and escalationSYSTEM_COMPONENTsharedGuide

A customer-support copilot assists a human agent with grounded answers, retrieval, drafting, summarization, and suggested actions while preserving escalation and approval boundaries.

Mental model

Start with assistance that keeps a human as decision owner, then automate only the support actions that have reliable knowledge, safe tools, and bounded consequences.

Why it matters

Support is a high-frequency environment where stale knowledge or an incorrect action can quickly affect real customers, accounts, and money.

Related concepts

Knowledge-base lifecycle

Read the full Guide
AI-Powered Solo BusinessExplore ai-powered solo business as part of the AI knowledge map.0

Prefer a clear path?

Use Courses for goal-oriented learning.

The same canonical knowledge is projected into 15 simpler learning paths, without duplicating the graph here.

Browse all courses
Advanced learning toolsOpen the legacy Guided Path and browser progressExpand this only when you need the older 10-stage Experience progression, review state, or Legacy Model Index.

GUIDED PATH · CURRENT EXPERIENCES

Guided Path

Ten stages from model behavior to production architecture. Stages show the full curriculum even when every mental model does not yet have its own page.

Progress is stored only in this browser for now. No account is required.

  1. STAGE 00

    AI Systems Mental Model

    Understand what the model can suggest and what the application must guarantee.

  2. STAGE 01

    Behavior, Prompt & Output Contracts

    Shape behavior while keeping authority, context and runtime responsibilities explicit.

    Mental models: 4

  3. STAGE 02

    Context, Retrieval & RAG

    Build evidence pipelines that stay relevant, authoritative and inspectable.

  4. STAGE 03

    Memory, Knowledge & Source of Truth

    Design memory as a lifecycle and authority problem, not a storage feature.

    Mental models: 5

  5. STAGE 04

    Tools, MCP & Capability Boundaries

    Give Agents capabilities without turning model intent into permission.

    Mental models: 6

  6. STAGE 05

    Agent Loop, State & Long-Running Work

    Control iteration, state, interruption and work that outlives one request.

  7. STAGE 06

    Reliability, Security & Human Control

    Bound retries, side effects, trust and human intervention under production pressure.

  8. STAGE 07

    Evaluation, Observability & Production Economics

    Use traces and evaluation evidence to decide what is safe and worthwhile to ship.

  9. STAGE 08

    Graphs, Delegation & Multi-Agent Systems

    Add orchestration only when decomposition creates more value than coordination cost.

  10. STAGE 09

    Production Architecture & Capstones

    Integrate the layers and own a defensible production release decision.

LEGACY MODEL INDEX · V0.9

Knowledge Map

Mental models are the durable units. Experiences are practice surfaces that can exercise several models at once.

00AI Systems Mental Model4
  • S00-M01Probabilistic behavior vs application guarantees
  • S00-M02Model claim vs runtime fact
  • S00-M03Structured output as an application contract
  • S00-M04Runtime enforcement vs model persuasion
01Behavior, Prompt & Output Contracts4
  • S01-M01Instruction authority and provenance
  • S01-M02Ambiguity, specificity and instruction conflict
  • S01-M03Examples shape behavior; they do not enforce policy
  • S01-M04Prompt vs Context vs Runtime responsibility boundary
02Context, Retrieval & RAG8
  • S02-M01Finite context budget
  • S02-M02Context structure and cache-friendly stability
  • S02-M03Compaction vs critical-information retention
  • S02-M04Evidence granularity and chunking
  • S02-M05Dense, sparse and hybrid retrieval
  • S02-M06Reranking as a second evidence-selection stage
  • S02-M07Freshness, authority and source priority
  • S02-M08Structured and metadata-aware retrieval
03Memory, Knowledge & Source of Truth5
  • S03-M01Working memory vs durable memory
  • S03-M02Memory write/read lifecycle and selection policy
  • S03-M03Memory vs authoritative source-of-truth conflict
  • S03-M04Expiry, invalidation, versioning and forgetting
  • S03-M05User memory vs domain knowledge boundary
04Tools, MCP & Capability Boundaries6
  • S04-M01Tool contract and schema design
  • S04-M02Tool-result validation and explicit error surfaces
  • S04-M03Capability boundary and least privilege
  • S04-M04Reversible vs irreversible actions
  • S04-M05Human approval at the high-risk boundary
  • S04-M06MCP protocol, application state, authorization and capability boundaries
05Agent Loop, State & Long-Running Work6
  • S05-M01Act → observe → verify loop
  • S05-M02Planning vs direct execution
  • S05-M03Bounded autonomy, termination and escalation
  • S05-M04Checkpoint, resumable state and recovery
  • S05-M05Interrupt and cancellation semantics
  • S05-M06Async and long-running task lifecycle
06Reliability, Security & Human Control5
  • S06-M01Timeout ambiguity: missing response is not confirmed failure
  • S06-M02Retry policy and retry amplification
  • S06-M03Idempotency boundary for repeated intent
  • S06-M04Compensation and recovery after side effects
  • S06-M05Trust boundary and defense in depth for untrusted context
07Evaluation, Observability & Production Economics7
  • S07-M01Traceability as causal execution history
  • S07-M02Observability as a diagnosis interface
  • S07-M03Evaluation environment and verifier design
  • S07-M04Outcome vs trajectory evaluation
  • S07-M05Dataset slices, regression and aggregate-improvement traps
  • S07-M06Confidence, variance and sample-size humility
  • S07-M07Cost, latency, quality and release vetoes as one decision
08Graphs, Delegation & Multi-Agent Systems4
  • S08-M01Loop vs Graph responsibility boundary
  • S08-M02Sequential, conditional and parallel topology trade-offs
  • S08-M03Delegation contract and shared vs isolated state
  • S08-M04Independent verification, coordination cost and correlated failure
09Production Architecture & Capstones3
  • S09-M01Cross-layer architecture decomposition
  • S09-M02Bounded rollout, fallback and graceful degradation
  • S09-M03Evidence synthesis into SHIP / BLOCK / INCONCLUSIVE