Library
Explore AI Engineering Topics
104+ in-depth guides on LLM agents, RAG pipelines, MCP, system design, and interview prep. Free topics are fully accessible — preview topics show a sample before sign-up.
LLM & Agentic
49LLM Lifecycle
PreviewComplete lifecycle of large language models from pre-training through fine-tuning, RLHF, and deployment — with architecture diagrams and production considerations.
How LLMs Call Tools
PreviewHow LLMs use function calling and tool use — the mechanics behind tool-calling agents, from prompt engineering to structured output.
Prompt Engineering Foundations
PreviewWhere an instruction belongs — tool description, system prompt, or user turn — why emphasis that worked on older models now causes over-triggering, and the prompt patterns that have expired.
Tool Surface Design
PreviewWhen an action deserves its own tool instead of bash, how to write a description that raises selection accuracy, and how to scale past a few dozen tools without destroying the prompt cache.
When to Fine-Tune: The Decision Framework
PreviewShould you fine-tune at all? A structured decision framework for prompt engineering vs RAG vs fine-tuning.
LoRA, QLoRA & PEFT Methods
PreviewHow LoRA works inside transformer layers, QLoRA for memory-efficient training, and the full PEFT method comparison with code examples.
Knowledge Distillation: Large to Small
PreviewTrain a small, fast model to mimic a large teacher — economics, pipeline, and quality filters for production distillation.
Fine-Tuning Data Preparation
PreviewData volume guidelines, quality checklists, and the complete preparation pipeline for fine-tuning datasets.
Cost, Latency & Quality Tradeoffs
PreviewThe economics of fine-tuning — quality/cost/latency triangle, real 2026 price points, break-even math, and the maintenance tax nobody budgets for.
Fine-Tuning Evaluation & Validation
PreviewHow to evaluate fine-tuned models — metrics by task type, regression testing, and the complete evaluation pipeline.
Preference Optimization
PreviewRLHF, DPO and RLAIF explained by what each one removes: how preference pairs teach judgement that demonstrations cannot, and why reward hacking and length bias are predictable.
Anatomy of a Tool Call
PreviewStep-by-step breakdown of an LLM tool call: request, schema validation, execution, and result handling with code examples.
Stateless vs Stateful
PreviewStateless vs stateful LLM architectures: trade-offs for agent design, conversation management, and production deployment.
The Agent Loop
FreeBuild a complete tool-calling AI agent in 15 lines of Python. Understand the core agent loop pattern that powers all LLM agents.
Agentic Spectrum
PreviewThe spectrum of AI agent architectures from simple prompt-response to fully autonomous multi-agent systems, with trade-offs at each level.
ReAct Pattern
FreeReAct (Reasoning + Acting) pattern for AI agents: how to combine chain-of-thought reasoning with tool use for better agent performance.
Reflexion
PreviewReflexion pattern for self-improving AI agents: verbal reinforcement learning where agents turn failures into written lessons and retry — no gradients required.
Hierarchical Delegation
PreviewHierarchical delegation pattern for multi-agent systems: orchestrating specialized agents with a coordinator for complex tasks.
Planner-Executor
PreviewPlanner-Executor agent pattern: separating planning and execution phases for more reliable and debuggable AI agent workflows.
Human in the Loop: Approval Gates
PreviewApproval gates for autonomous agents — placing them by reversibility, the confirmation round trip and the deadlocks it hides, denials that redirect rather than stall, and what to do when no human is watching.
Agent Durability
PreviewMake a long agent run survive a crash: the window where a side effect lands before anything records it, idempotency keys that survive replay, and what to checkpoint.
The Four Types of Agent Memory
PreviewWorking, short-term, long-term, and org-wide memory — what each type stores, its lifetime, and implementation.
Memory Implementation & Patterns
PreviewProduction-ready memory code, episodic vs semantic vs procedural memory, and system design patterns.
Memory Decisions & Worked Examples
PreviewWhen to use which memory type, with three real-world scenarios: customer support, sales agent, and code assistant.
Memory Compliance & Interview Prep
PreviewGDPR and privacy compliance for agent memory, plus FDE interview scenarios and deep-dive questions.
Messages API
PreviewClaude Messages API deep dive: request/response format, system prompts, multi-turn conversations, and best practices.
Tool Use with Claude
PreviewImplement tool use with Claude API: define tools, handle tool calls, and build reliable function-calling agents.
Streaming with Claude
PreviewStream Claude API responses for real-time UX: server-sent events, token-by-token rendering, and production streaming patterns.
Structured Output
PreviewGet structured JSON output from Claude: constrained generation, schema validation, and reliable data extraction patterns.
Prompt Caching
PreviewCut Claude cost and latency by up to 90% with prompt caching: cache the stable prefix (system prompt, tools, documents, history), pay full price once, then read at ~10%.
Adaptive Thinking & Effort
PreviewLet Claude reason before it answers: adaptive thinking, the five effort levels, why budget_tokens is gone, summarized vs omitted thinking blocks, and preserving them across tool calls.
Agent SDK Patterns
PreviewProduction patterns for building AI agents with the Claude Agent SDK: the loop, subagents, permissions, context management, MCP, and hooks — plus when to use the SDK vs the raw API.
Who Owns the Loop
PreviewManual loop vs the SDK Tool Runner vs Managed Agents vs the Claude Agent SDK — separated by who supplies the agent harness and who supplies the deployment.
Context Management
PreviewContext editing, compaction and memory as Claude API features: what each does to the conversation, and why the compaction block must be appended back verbatim.
Agent Cost Control
PreviewControl what an agent costs: effort levels, task budgets vs session budgets vs max_tokens, prompt-cache economics, model routing, and per-turn token accounting.
Metrics Framework
PreviewObservability metrics for LLM applications: latency, token usage, cost tracking, and quality scoring dashboards.
Eval & Observability
PreviewComplete guide to LLM evaluation and observability: automated evals, human feedback loops, A/B testing, and monitoring.
Agentic Evaluation
PreviewEvaluate AI agents in production: task completion metrics, trajectory analysis, and automated agent quality benchmarks.
Your First Agent SDK Agent
PreviewRun a Claude Agent SDK agent in ten lines: query() vs ClaudeSDKClient, every message type the stream yields, and configuring model, budget and tool restrictions through ClaudeAgentOptions.
Custom Tools: @tool and SDK MCP Servers
PreviewBuild custom Claude Agent SDK tools with @tool and create_sdk_mcp_server, get the mcp__server__tool naming right, and see four production tools — SQL, HTTP, filesystem, shell — with the guardrail each needs.
Permissions & Hooks: Gating the Agent
PreviewGate a Claude agent: permission modes, disallowed_tools vs allowed_tools, deciding per call with can_use_tool, and using PreToolUse and PostToolUse hooks for audit logging and redaction.
Worked Example: Repo Review Agent
PreviewA complete Claude Agent SDK agent: a custom diff tool, a stripped tool surface, a per-call path gate, an audit hook, a reviewer subagent, and a budget cap — structurally unable to modify what it reviews.
Why a Graph? State, Nodes & Edges
PreviewBuild a LangGraph agent from scratch: StateGraph, TypedDict state, the add_messages reducer, and the decision table for when a graph beats a plain while-loop.
Tools & Routing
PreviewDefine LangGraph tools with @tool, wire ToolNode and tools_condition into an agent loop, use the prebuilt create_agent, and write custom routers that branch on your own state.
Checkpointers, Threads & Store
PreviewLangGraph persistence end to end: checkpointers and thread_id for resumable runs, get_state and time travel, PostgresSaver in production, and the Store for memory that outlives a thread.
interrupt(): Approval Gates
PreviewPause a LangGraph run for human review with interrupt(), resume it with Command, and build one gate that supports approve, edit and reject without double-charging on re-run.
Subgraphs & the Supervisor
PreviewMulti-agent LangGraph: routing and updating state in one Command return, using a compiled graph as a node, and choosing between shared and isolated history for a worker.
Worked Example: Support Triage Agent
PreviewA complete runnable LangGraph agent in sixty lines: two tools, a router that knows which is dangerous, an approval gate on refunds, and a checkpointer so the pause survives a restart.
Productionising LangGraph
PreviewDeploy a LangGraph agent: API worker vs queue consumer vs scheduled job on one Postgres, stream modes per consumer, bounding steps and tokens, and testing routers without a model.
RAG & MCP
20RAG Architecture
PreviewRetrieval-Augmented Generation (RAG) architecture explained: ingestion pipeline, vector search, prompt augmentation, and production patterns.
Document Processing
PreviewDocument processing for RAG pipelines: SFTP and connector ingestion, PDF parsing, OCR, table extraction, and multi-modal document understanding.
Chunking Strategies
PreviewText chunking strategies for RAG: fixed-size, semantic, recursive, and document-aware chunking with performance comparisons.
Embedding & Indexing
PreviewEmbedding models and vector indexing for RAG: choosing embeddings, HNSW vs IVF, dimensionality, and index optimization.
Embeddings
PreviewUnderstanding embeddings for AI applications: text, image, and multi-modal embeddings with similarity search and clustering.
Metadata Strategies
PreviewMetadata strategies for RAG: filtering, hybrid search, metadata extraction, and structured metadata for improved retrieval.
Retrieval & Reranking
PreviewAdvanced retrieval and reranking for RAG: BM25, dense retrieval, cross-encoder reranking, and hybrid search strategies.
RAG Evaluation
PreviewEvaluate RAG system quality: retrieval precision/recall, answer faithfulness, and end-to-end pipeline benchmarking.
RAGAS Framework
PreviewRAGAS evaluation framework for RAG: faithfulness, answer relevancy, context precision, and automated quality scoring.
RAG Monitoring
PreviewProduction monitoring for RAG systems: retrieval quality dashboards, drift detection, and automated alerting.
Advanced RAG Patterns
PreviewAdvanced RAG techniques: query decomposition, self-RAG, corrective RAG, adaptive retrieval, and multi-hop reasoning.
RAG Best Practices
PreviewProduction RAG best practices: pipeline optimization, failure handling, testing strategies, and common pitfalls to avoid.
MCP Overview
FreeModel Context Protocol (MCP) explained: the open standard for connecting AI models to tools, data sources, and external systems.
MCP Architecture
PreviewMCP architecture deep dive: client-server model, protocol layers, message types, and connection lifecycle.
Building MCP Servers
PreviewBuild MCP servers step-by-step: Python and TypeScript implementations with tools, resources, and prompts.
MCP Transport
PreviewMCP transport layers: stdio, SSE, and streamable HTTP transports with implementation details and trade-offs.
MCP Discovery
PreviewMCP tool discovery and capability negotiation: how clients discover server capabilities and tools dynamically.
MCP Security
PreviewSecurity considerations for MCP: authentication, authorization, input validation, and sandboxing strategies.
MCP in Production
PreviewDeploy MCP servers in production: scaling, monitoring, error handling, and reliability patterns.
MCP on GCP
PreviewRun MCP on Google Cloud Platform: Cloud Run deployment, IAM integration, and GCP-native tool implementations.
System Design
32System Design 101
FreeSystem design fundamentals for AI engineers: client-server, APIs, latency percentiles, caching, load balancing, databases, and queues — each explained from zero, then mapped to how LLM systems change it.
AI System Design Vocabulary
FreeThe 60-term plain-English glossary for AI system design: LLM basics, retrieval, agents, infrastructure, reliability, scaling, cost, and safety — with deep-dive links into every guide section.
Your First Agentic System
FreeBuild a support bot end to end: six iterations from one API call to a production-shaped architecture with retrieval, caching, model routing, guardrails, and observability — runnable code at every step.
The Paradigm Shift
FreeTraditional vs agentic system design: the 7 dimensions that transform, anatomy of an agentic system, control flow paradigms, failure modes, and when to go agentic.
5-Phase Framework
FreeFive-phase system design framework for AI interviews: requirements, architecture, data flow, scaling, and production readiness.
10-Layer Architecture
PreviewStaff-level 10-layer architecture for AI-native systems: from infrastructure to user experience, with production examples.
Scaling 10k to 1M
PreviewScale AI systems from 10K to 1M users: caching, sharding, async processing, and infrastructure evolution strategies.
Real-World Case Studies
PreviewHow OpenAI Deep Research, Claude Code, Perplexity, Cursor, and Devin work under the hood. Production architecture breakdowns with design decisions for interviews.
Security Overview
PreviewSecurity and privacy for AI applications: threat models, data protection, compliance frameworks, and defense-in-depth.
Guardrails & Safety
PreviewAI guardrails and safety: content filtering, output validation, safety classifiers, and responsible AI deployment.
PII Detection
PreviewPII detection and redaction in LLM applications: entity recognition, masking strategies, and compliance automation.
Prompt Injection Defense
PreviewDefend against prompt injection attacks: detection techniques, input sanitization, and multi-layer defense strategies.
Multi-Tenant Isolation
PreviewMulti-tenant isolation for AI platforms: data separation, model isolation, rate limiting, and tenant-aware architectures.
Audit & Compliance
PreviewAudit logging and compliance for AI systems: SOC2, HIPAA, GDPR requirements, and automated compliance monitoring.
Context Engineering
PreviewThe discipline replacing prompt engineering: designing dynamic context systems that give agents the right information at the right time. Covers lazy loading, memory-augmented architectures, and production patterns.
Semantic Caching
PreviewSemantic caching for LLM applications: reduce costs and latency by caching semantically similar queries with vector similarity.
Model Routing
PreviewTiered model routing: route queries to the right model (GPT-4, Claude, Haiku) based on complexity, cost, and latency requirements.
Rate Limiting & Cost
PreviewRate limiting and cost management for LLM APIs: token budgets, per-user quotas, and cost optimization strategies.
Inference Optimization
PreviewLLM inference optimization: batching, quantization, KV-cache, speculative decoding, and hardware selection.
Hallucination Detection
PreviewDetect and prevent LLM hallucinations: factuality checking, grounding verification, and off-brand content filtering.
Agent Failure Modes
PreviewCommon AI agent failure modes: infinite loops, tool misuse, context window overflow, and recovery strategies.
Deployment & Rollout
PreviewDeploy and roll out AI systems: canary releases, feature flags, A/B testing, and safe rollback strategies.
Event-Driven Async
PreviewEvent-driven async architectures for AI: message queues, webhook patterns, and asynchronous agent orchestration.
Data Flywheel
PreviewBuild a data flywheel for AI products: feedback loops, continuous learning, and data-driven model improvement cycles.
Deploy Your First Agent
FreeTake an agent from a script on your laptop to production, step by step: HTTP service, Docker, secrets, the four runtime shapes, state schema, environments, a CI pipeline with an eval gate, and the first 24 hours live. Explained for non-engineers and staff engineers at once.
Agent Tech Stack
PreviewThe ten layers of a production agent stack — model access, gateway, orchestration, runtime, state, vector, queue, observability, evals, secrets — with the choice criterion for each, three fully costed reference stacks, and the four buy-vs-build calls that matter.
Multi-Agent in Production
PreviewRunning multi-agent systems as deployed infrastructure: four runtime topologies compared, the task envelope contract between agents, fleet versioning, distributed tracing across hops, retry-storm math, and a four-agent contract-review pipeline built end to end.
Scaling the Agent Runtime
PreviewWhy CPU autoscaling is the wrong signal for a 40-second agent run: in-flight-run scaling, queueing-theory capacity math, the derived timeout ladder, connection pooling, a five-rung degradation ladder, token-rate limiting, and a load-test recipe.
Day-2 Operations
PreviewKeeping an agent alive after launch: the six drift sources, SLOs that measure the product instead of the server, the maintenance calendar, incident response for silent quality failures, runbooks you can run at 3am, the model-deprecation migration, and cost governance.
Architecture Decision Records
PreviewADRs for AI systems: a reusable template plus three fully worked records (model provider choice, build-vs-buy guardrails, sync-to-async migration) — each with an explicit revisit trigger.
Migrations & Rollouts
PreviewStaff-level AI migration playbooks: provider swaps, monolith-to-agentic decomposition, model-version upgrades, and zero-downtime component swaps under one safe shadow → canary → ramp → rollback pattern.
Org & Leadership
PreviewThe staff organizational layer: platform-vs-product ownership, incident command for silent LLM-quality failures, cost governance, the RFC process, and driving adoption without authority.
Interview Prep
3Cost vs Latency
PreviewLLM inference cost vs latency trade-offs: optimization strategies for production AI systems with budget constraints.
Code Leakage Prevention
PreviewPrevent code and data leakage in LLM applications: sandboxing, output filtering, and secure coding practices.
The Interview Playbook
PreviewOne end-to-end interview playbook for AI engineers: the coding bar, a six-step process, the seven discovery dimensions, communication structure, scripts to say out loud, recovery lines, and four worked examples.
Full Access
Unlock All 74+ Sections
Get full access to the complete guide including worked problems, interview scripts, checklists, and staff-level deep dives.
Unlock Premium