AI Engineering Glossary

Plain-English definitions for agentic AI, orchestration, and AI engineering terms. Whether you're building your first project using AI agent or scaling an existing one, find every term you need in one place.

101 terms found

A

  • A2A Protocol

    The A2A protocol, short for Agent2Agent, is an open standard for communication between AI agents built by different teams on different frameworks. It lets one agent discover another's capabilities, delegate a task to it, and track that task to completion, without either side exposing its internal prompts, models, or tools.

  • Agent Memory

    Agent memory is the set of mechanisms that let an AI agent retain and reuse information beyond a single context window: session state, notes written to files, databases, and vector stores. It is what allows an agent to carry facts, preferences, and lessons across tasks instead of starting every session from zero. The gap it fills is well documented: on the LoCoMo benchmark of very long conversations, averaging 300 turns across up to 35 sessions, Maharana and colleagues found in 2024 that LLM agents substantially lag human performance at recalling earlier context.

  • Agent Orchestration

    Agent orchestration is the coordination of multiple AI agents working on parts of a larger task: assigning work, sequencing steps, passing context between agents, and merging results into one output. The coordinator can be a lead agent that decides dynamically or a deterministic controller that follows fixed routing rules.

  • Agent Swarm

    An agent swarm is a group of AI agents launched in parallel on the same problem or on many slices of it, with their outputs merged, ranked, or voted on afterward. The approach trades extra compute and tokens for broader coverage and higher confidence than any single agent run can provide.

  • Agentic AI

    Agentic AI is a class of AI systems that pursue a goal by planning steps, taking actions through tools, observing the results, and adjusting course, all with limited human direction. Unlike a model that returns one answer to one prompt, an agentic system keeps working until the goal is met or a stop condition triggers. The ceiling on that work keeps rising: METR measured in 2025 that the length of tasks AI agents can complete autonomously has been doubling roughly every 7 months for the last six years.

  • Agentic Automation

    Agentic automation is business process automation in which AI agents handle the steps that require judgment, such as interpreting messy inputs, choosing among options, and recovering from exceptions. Traditional rule-based automation executes fixed scripts and breaks when inputs vary; agentic automation adapts, escalating to people only when a decision exceeds its mandate. Adoption has moved past pilots: 46% of leaders in Microsoft's 2025 Work Trend Index said their organization uses AI agents to fully automate workstreams or business processes, led by customer service, marketing, and product development.

  • Agentic Coding

    Agentic coding is the practice of producing software by directing an AI coding agent that reads the codebase, edits files, runs commands, and iterates on the results, instead of typing the code yourself. The developer supplies intent, constraints, and review; the agent supplies the keystrokes. The practice already operates at real scale: GitHub's Octoverse 2025 report counted over 1 million pull requests generated by its Copilot coding agent between May and September 2025, concentrated in established repositories rather than experimental projects.

  • Agentic Design Patterns

    Agentic design patterns are reusable structures for building AI agent systems. The core four are reflection, tool use, planning, and multi-agent collaboration. Each pattern describes a proven way to arrange model calls, tools, and control flow, and each trades off autonomy, cost, latency, and reliability in a different, predictable way.

  • Agentic Engineering

    Agentic engineering is the discipline of directing AI agents to build software through specifications, guardrails, and review instead of writing every line by hand. The engineer's job shifts from typing code to defining what correct looks like, constraining what the agent can do, and verifying what it produced. The raw material for this shift is already everywhere: the 2025 Stack Overflow survey found 84% of developers using or planning to use AI tools in their workflow, up from 76% the year before.

  • Agentic RAG

    Agentic RAG is retrieval-augmented generation in which an AI agent controls the retrieval process. Instead of running one fixed search before answering, the agent decides whether to retrieve at all, chooses among sources, reformulates queries, evaluates what came back, and searches again until it has enough evidence to answer well.

  • Agentic Workflow

    An agentic workflow is a multi-step process in which AI agents plan, take actions, evaluate results, and hand off work toward a defined outcome. It combines model-driven decisions with deterministic steps such as scripts, API calls, and approval gates, so the system can adapt to what it finds while still producing predictable, verifiable output. The ceiling on what these workflows can handle keeps rising: METR measured in 2025 that the length of tasks frontier agents can complete at a 50% success rate has doubled roughly every 7 months since 2019, with Claude 3.7 Sonnet reaching a time horizon of about 50 minutes.

  • AGENTS.md

    AGENTS.md is a plain markdown file placed in a repository to tell AI coding agents how the project works: build and test commands, code conventions, architecture notes, and rules to follow. It is a vendor-neutral convention, described as a README for agents, and it is read automatically by most coding agents as of 2026.

  • AI Agent

    An AI agent is a software system that uses a language model to decide what to do next, then acts through tools such as APIs, files, shells, or a browser. It observes the result of each action and keeps iterating until the goal is complete, failed, or handed back to a person. Demand for these systems is broad: 82% of business leaders in Microsoft's 2025 Work Trend Index expect to use digital labor such as AI agents to expand their workforce within 12 to 18 months.

  • AI Agent Architecture

    AI agent architecture is the structural design of an agent system: the language model at its core, the tool layer it acts through, the memory it persists, the loop that drives planning and execution, and the guardrails that bound its behavior, plus the connections that turn those parts into one working system. Getting this structure right is where the field's progress is coming from: Stanford's 2025 AI Index recorded SWE-bench scores rising 67.3 percentage points in a single year as agentic systems matured.

  • AI Agent Evaluation

    AI agent evaluation is the practice of measuring whether an AI agent completes tasks correctly end to end, scoring the full trajectory of decisions, tool calls, and intermediate steps as well as the final outcome. It extends single-response evaluation to multi-step autonomous work, where an agent can reach a right answer through a broken process or fail ten steps after a good start. The stakes keep rising because agents keep getting more capable: METR measured in 2025 that frontier models like Claude 3.7 Sonnet have a 50 percent task-completion time horizon of about 50 minutes, a horizon that has been doubling roughly every seven months since 2019.

  • AI Agent Framework

    An AI agent framework is a software library that supplies the building blocks for creating AI agents: model calls, tool definitions, memory, state management, and the loop that lets an agent plan, act, and observe results. Developers assemble these components instead of wiring the agent loop from scratch for every project. The audience for this tooling is already large: in LangChain's 2024 State of AI Agents survey of more than 1,300 professionals, 51% had agents in production, rising to 63% at mid-sized companies of 100 to 2,000 employees.

  • AI Agent Security

    AI agent security is the practice of protecting systems from the risks created when AI agents act autonomously: following injected instructions hidden in data, using over-broad permissions, installing hallucinated dependencies, or pushing unreviewed changes to production. It combines classic access control with defenses specific to how language models can be manipulated.

  • AI Alignment

    AI alignment is the research field concerned with making AI systems reliably pursue the goals their operators actually intend, rather than proxies, loopholes, or literal readings of an instruction. As models grow more capable and act with more autonomy, the gap between what was asked and what was meant becomes a safety and reliability problem.

  • AI Code Documentation

    AI code documentation is the practice of generating and maintaining documentation, including comments, READMEs, API references, and architecture notes, with AI models that read the source directly. Because the docs are derived from the code itself, they can be regenerated whenever the code changes, which attacks the oldest problem in documentation: drift between what the docs say and what the code does. Developers have embraced this faster than almost any other AI use: in the 2025 Stack Overflow Developer Survey, 30.8% said they mostly use AI for documenting code and another 30.3% use it partially.

  • AI Code Generation

    AI code generation is the production of source code by a machine learning model from natural-language descriptions, existing code, or both. It spans everything from single-line autocomplete to entire applications built by autonomous agents, and it has become the default first draft mechanism for a large share of professional software in 2026. The demand shows up in model usage itself: the Anthropic Economic Index found that 37.2% of queries sent to Claude fell into the computer and mathematical category, covering tasks like software modification and debugging, making coding the single largest use of the model.

  • AI Code Refactoring

    AI code refactoring is the use of AI coding agents to restructure existing code without changing what it does: renaming, extracting functions, splitting oversized modules, and migrating deprecated patterns across a codebase. The test suite defines the behavior that must survive, the agent performs the mechanical edits, and a human reviews the result before it merges.

  • AI Code Review

    AI code review is the use of a large language model to examine code changes for bugs, security issues, logic errors, and style problems, typically as an automated first pass on pull requests. It supplements human review by catching mechanical defects early, leaving people to judge architecture, intent, and fitness for purpose.

  • AI Coding Agent

    An AI coding agent is software that takes a development task in natural language, then reads the codebase, writes and edits files, runs commands, and iterates on the results until the task is done. It operates with far more autonomy than editor autocomplete, completing whole units of work rather than suggesting the next line. The category matured fast: when the SWE-bench benchmark launched in 2023, the best model resolved only 1.96% of real GitHub issues drawn from 12 popular Python repositories.

  • AI Coding Assistant

    An AI coding assistant is a tool that helps developers write software by generating suggestions, completing code, and answering questions inside the editor, powered by a large language model. The developer stays in control of every change, which separates assistants from autonomous agents that edit files and run commands on their own. Adoption is broad: in JetBrains' 2025 survey of 24,534 developers across 194 countries, 85% regularly used AI tools for coding and 62% relied on at least one AI coding assistant, agent, or AI code editor.

  • AI DevOps

    AI DevOps is the application of AI agents to the software delivery pipeline: diagnosing CI failures, reviewing infrastructure changes, managing deploys, and running the first pass of incident response. Software that can read logs, query dashboards, and execute commands absorbs the operational toil that used to interrupt engineers, while humans keep approval over anything that touches production. The shift is already mainstream: 75% of respondents to the 2024 DORA survey said they rely on AI for at least one daily professional responsibility, and in the 2025 DORA report, based on nearly 5,000 technology professionals, 90% reported using AI at work with more than 80% saying it increased their productivity.

  • AI Governance

    AI governance is the set of policies, roles, and controls an organization uses to deploy AI responsibly: deciding which uses are allowed, who is accountable for each system, how data and risk are managed, and how compliance is demonstrated. It turns scattered AI adoption into something the organization can actually see, steer, and defend. The need is growing: reported AI-related incidents rose to a record 233 in 2024, a 56.4% increase over 2023, according to the AI Incidents Database figures in Stanford's AI Index.

  • AI Grounding

    AI grounding is the practice of tying a model's output to verifiable sources such as retrieved documents, database records, or tool results, so every claim can be traced and checked. A grounded answer cites where its facts came from; an ungrounded answer relies only on what the model absorbed during training.

  • AI Guardrails

    AI guardrails are the technical controls placed around an AI system to keep its behavior inside acceptable bounds: input and output filters, restricted permissions, validation checks, sandboxed execution, and mandatory approval gates before consequential actions. They assume the model will sometimes be wrong or manipulated, and limit the damage when it is.

  • AI Hallucination

    An AI hallucination is confident output from a model that is factually wrong or entirely invented: a function that does not exist in the library, a parameter the API never accepted, a package nobody published, a citation to a paper never written. The output is fluent and plausible, which is exactly what makes it dangerous. The problem is measurable at scale: a 2024 Stanford study found LLMs hallucinated on 58 percent (ChatGPT) to 88 percent (Llama 2) of verifiable questions about random federal court cases.

  • AI IDE

    An AI IDE is a development environment designed around AI capabilities from the ground up, with codebase-aware chat, predictive multi-file editing, and built-in agent modes as core features rather than plugins. The editor treats the model as a first-class collaborator with access to the same project the developer sees. The broader category is now mainstream: 62% of developers rely on at least one AI coding assistant, agent, or AI code editor, per JetBrains' 2025 survey of 24,534 developers.

  • AI Inference

    AI inference is running a trained model to produce output: the computation that happens every time an application sends a prompt and gets a response. Training builds the model once; inference is the recurring workload after that, and it is what API pricing, latency budgets, and GPU serving infrastructure all revolve around. That workload is growing at a startling pace: Google reported its monthly inference volume grew 50x in one year, from 9.7 trillion tokens a month to over 480 trillion by May 2025.

  • AI Observability

    AI observability is the practice of instrumenting AI systems in production, capturing traces, prompts, tool calls, token costs, latency, and quality signals, so teams can see what a model or agent actually did and why. It extends traditional observability to systems whose behavior is probabilistic and whose failures rarely throw exceptions. The industry has largely accepted the argument: nearly 89 percent of organizations building agents have implemented observability for them, per LangChain's late-2025 survey of 1,300+ professionals.

  • AI Pair Programming

    AI pair programming is working with an AI model as a continuous collaborator while writing software: discussing approaches, generating and critiquing code, and reviewing decisions as they happen. It mirrors the driver and navigator roles of human pair programming, with the AI able to play either seat depending on the task.

  • AI Red Teaming

    AI red teaming is the practice of deliberately attacking your own AI system, probing it with jailbreaks, injection payloads, and misuse scenarios, to find failures before real adversaries or users do. It adapts the adversarial mindset of security red teams to model behavior, covering both malicious attacks and ordinary inputs that produce harmful output.

  • AI Sandbox

    An AI sandbox is an isolated execution environment where an AI agent can run code, install packages, and modify files without touching production systems, real credentials, or the host machine. Anything the agent breaks stays inside the boundary, which makes autonomous work safe enough to delegate at scale.

  • AI Slop

    AI slop is low-quality, high-volume content generated by AI and published or shipped without meaningful human judgment. The term covers articles, images, code, and pull requests alike: output that is fluent enough to pass a glance but adds noise instead of value, the visible symptom of automation running without a review process behind it.

  • AI Sycophancy

    AI sycophancy is a language model's tendency to tell users what they want to hear instead of what is true: agreeing with stated opinions, validating flawed plans, and reversing correct answers under pushback. It emerges from preference training, where models learn that agreeable responses earn higher human ratings than accurate ones.

  • AI Testing

    AI testing covers two related practices: using AI models to generate, run, and maintain software tests, and testing AI systems themselves, whose outputs vary from run to run and resist traditional pass-fail assertions. Both aim at the same goal, catching defects before users do, but each demands different techniques because deterministic and probabilistic software fail in different ways.

  • AI-Assisted Coding

    AI-assisted coding is a way of working where a developer writes software with continuous help from AI, through autocomplete, inline edits, and chat, while remaining the author of every change. The human drives; the model accelerates. This distinguishes it from agentic coding, where an agent executes whole tasks on its own. The practice is now the norm rather than the exception: in the 2025 Stack Overflow survey of more than 49,000 developers, 84% were using or planning to use AI tools, up from 76% in 2024, even as the share who distrust the accuracy of AI output rose from 31% to 46%.

  • Attention Mechanism

    An attention mechanism is the component of a transformer model that computes, for each token, how much every other token in the sequence should influence it. These learned relevance scores are how language models resolve references, track structure, and use context, and their cost is what makes long prompts expensive. The idea proved powerful enough to stand alone: the 2017 paper "Attention Is All You Need" dropped recurrence and convolutions entirely and still set a single-model state of the art of 41.8 BLEU on WMT 2014 English-to-French translation.

  • Automation Bias

    Automation bias is the human tendency to trust output from an automated system more than the evidence warrants, accepting a machine's answer with less scrutiny than the same claim would get from a person. In AI-assisted engineering it appears as approving agent-written code, configs, and analyses largely because they arrive looking complete, confident, and professionally formatted. Developer trust is poorly calibrated in both directions: in the 2025 Stack Overflow survey only 33% of developers trusted AI output accuracy, with just 3% highly trusting it and 46% actively distrusting it, yet 51% of professional developers used AI tools daily.

  • Autonomous AI Agent

    An autonomous AI agent is an AI agent that completes multi-step goals end to end without a human approving each action. Safety comes from constraints set in advance, such as scoped permissions, budgets, and guardrails, plus review of the finished result, rather than from supervision during the run. How much an agent can safely be handed keeps growing: METR found in 2025 that the length of tasks frontier agents can complete with 50% reliability has doubled roughly every 7 months for six years.

B

  • Browser Agent

    A browser agent is an AI agent whose primary tool is a web browser. Given a goal, it navigates pages, reads content, clicks elements, fills forms, and completes multi-step flows the way a person would, letting it operate any website, including the vast majority that expose no API. Reliability on the open web remains the hard part: on WebArena's realistic tasks spanning e-commerce, forums, code hosting, and CMS work, the best GPT-4 agent in the 2023 Carnegie Mellon study completed 14.41% of tasks end to end, compared with 78.24% for humans.

C

  • Chain-of-Thought Prompting

    Chain-of-thought prompting is the technique of instructing a language model to work through a problem step by step before giving its final answer, rather than answering immediately. Making the intermediate reasoning explicit measurably improves accuracy on math, logic, planning, and multi-step tasks, and it leaves a visible trail you can inspect when the answer is wrong. In the original 2022 paper from Wei et al. at Google, a 540B-parameter model prompted with just eight chain-of-thought exemplars reached state-of-the-art accuracy on the GSM8K math benchmark, beating even a finetuned GPT-3 with a verifier.

  • CLAUDE.md

    CLAUDE.md is the markdown instruction file that Claude Code, Anthropic's coding agent, automatically reads at the start of every session. It carries persistent project context: build and test commands, code conventions, architecture notes, and rules the agent must follow, so the same guidance applies to every task without being repeated in prompts.

  • Code LLM

    A code LLM is a large language model trained or fine-tuned primarily on source code, optimized to generate, complete, explain, and repair programs. The term covers dedicated code models and, increasingly, general frontier models whose training was weighted heavily toward code because programming became their most commercially important skill.

  • Comprehension Debt

    Comprehension debt is the accumulated body of code in a production system that nobody on the team genuinely understands, created when AI-generated changes ship without real review. The code works and the tests pass, but the knowledge that normally forms in a human head by writing the code never formed anywhere, and the gap compounds with every unexamined merge.

  • Computer Use

    Computer use is an AI agent capability where the model operates a computer through its graphical interface: it looks at screenshots, moves the cursor, clicks, types, and scrolls, exactly as a person would. This lets an agent work with any application on screen, including software that offers no API for structured access.

  • Context Engineering

    Context engineering is the discipline of controlling everything an AI model sees at inference time: which files, documents, conversation history, instructions, and tool results enter the context window, in what form, and in what order. Where prompt engineering shapes a single instruction, context engineering manages the model's entire field of view across a long-running task.

  • Context Rot

    Context rot is the gradual degradation of a language model's output quality as its context window fills with stale, contradictory, or irrelevant information over a long session. The model starts confusing old instructions with current ones, repeating abandoned approaches, and losing track of what actually matters, even though nothing about the model itself has changed.

  • Context Window

    A context window is the maximum amount of text an AI model can process in a single request, measured in tokens, spanning the system prompt, conversation history, retrieved documents, tool results, and the model's own output. It functions as the model's working memory: anything inside it can influence the answer, and anything outside it does not exist for the model.

D

  • Data Poisoning

    Data poisoning is an attack that corrupts the data an AI system learns from or retrieves, so the model absorbs attacker-chosen behavior: hidden backdoors, degraded accuracy, or planted misinformation. It is a supply-chain attack on the data layer, and it can target pretraining corpora, fine-tuning sets, or the documents a RAG pipeline trusts.

F

  • Few-Shot Prompting

    Few-shot prompting is the technique of including a small number of worked input-output examples directly in the prompt so the model infers the pattern you want and applies it to new input. It teaches by demonstration instead of description, and no training or fine-tuning is involved: the examples live in the prompt itself.

  • Foundation Model

    A foundation model is a large AI model trained once on broad data at scale, then adapted to many downstream tasks through prompting, fine-tuning, or tool integration. Instead of building a separate model per task, teams build products on a shared base, which is the economic pattern underlying the entire modern AI industry.

  • Frontier Model

    A frontier model is an AI model at the current edge of capability: one of the small set of systems that defines the state of the art at any moment. The term marks a moving boundary rather than a fixed spec, and it doubles as a regulatory category, since safety frameworks and compute-threshold rules target exactly this class.

H

  • Human in the Loop

    Human in the loop is a workflow design in which an AI system does the work but a person reviews and approves the decisions that matter before they take effect. The AI handles volume and speed; the human supplies judgment, accountability, and a veto at the points where errors are expensive or irreversible. The pattern persists partly because trust has limits: the 2025 DORA survey found 30% of technology professionals report little or no trust in AI-generated code, a key reason human review remains a standard checkpoint.

K

  • KV Cache

    The KV cache is the memory where a transformer language model stores the attention keys and values it has already computed for previous tokens. By reusing them, the model generates each new token without reprocessing the entire input, which is what makes token-by-token generation fast and what makes long contexts expensive in memory.

L

  • Large Language Model (LLM)

    A large language model (LLM) is a neural network trained on enormous amounts of text to predict the next token in a sequence. At sufficient scale, that single objective produces broad ability to write, summarize, reason, and generate code, which is why LLMs power chat assistants, coding agents, and most modern AI products. Adoption reflects that reach: in the 2025 Stack Overflow survey, 84% of developers said they use or plan to use AI tools in their work, up from 76% the year before.

  • LLM Benchmark

    An LLM benchmark is a standardized test suite used to compare language models on a shared task, coding, reasoning, knowledge, math, or agentic work, under fixed conditions so scores are comparable across models. Benchmarks are useful for shortlisting candidates and tracking industry progress, and increasingly compromised by models training on the test material.

  • LLM Evals

    LLM evals are repeatable, automated tests that measure whether a language model's output meets a defined quality bar: correctness, tone, safety, format, task completion. Where unit tests check that code behaves deterministically, evals score probabilistic output against a standard, so teams can change prompts and models without guessing whether quality moved.

  • LLM Fine-Tuning

    LLM fine-tuning is the practice of continuing a pretrained language model's training on a smaller, task-specific dataset so its default behavior shifts toward that task. Unlike prompting or retrieval, which steer the model at request time, fine-tuning changes the weights themselves, producing a specialized variant of the base model.

  • LLM Jailbreak

    An LLM jailbreak is an input crafted to make a language model ignore its safety training and produce output it was built to refuse. Techniques range from fictional role-play framings to encoded payloads and long multi-turn manipulation, and they matter to builders because every deployed model inherits the jailbreak surface of its base model.

  • LLM Quantization

    LLM quantization is the process of storing a language model's weights at lower numeric precision, for example 4-bit integers instead of 16-bit floats, so the model needs less memory and runs faster. Done well, it cuts hardware requirements by half to three quarters while losing only a small amount of output quality.

  • LLM Streaming

    LLM streaming is the delivery of a language model's response incrementally, token by token, as it is generated, instead of waiting for the full completion. It turns a response that takes many seconds to finish into one that starts appearing almost immediately, which is why nearly every chat and agent interface uses it.

  • LLM Temperature

    LLM temperature is a sampling parameter that controls how random a language model's output is. At low temperature the model almost always picks its highest-probability next token, giving consistent, predictable responses. At high temperature it samples more freely from less likely tokens, giving varied and sometimes surprising output.

  • LLM Token

    An LLM token is the unit a language model reads and writes text in: a short chunk of characters, on average about three-quarters of an English word. Tokens are also the billing unit for every model API, so token counts determine both what fits in a request and what each request costs.

  • LLM-as-a-Judge

    LLM-as-a-judge is an evaluation technique where one language model grades another model's output against a written rubric, scoring qualities like correctness, helpfulness, or faithfulness to a source. It makes qualitative evaluation cheap and fast enough to run on every change, at the cost of inheriting the judge model's own biases and blind spots. The technique earned its credibility in 2023, when Zheng et al. showed GPT-4 as a judge matched human preferences with over 80 percent agreement on MT-Bench and Chatbot Arena, the same level of agreement humans reach with each other.

  • llms.txt

    llms.txt is a proposed web standard: a markdown file at a site's root path (/llms.txt) that gives large language models a curated map of the site's most important content, with links to clean, markdown-friendly versions of key pages. It exists because HTML pages full of navigation and scripts waste an AI system's limited context.

M

  • Mixture of Experts

    Mixture of experts (MoE) is a neural network architecture in which each token activates only a small subset of the model's parameters, chosen by a learned router. The model can hold a very large total parameter count while spending the compute of a much smaller one on every request, which lowers inference cost.

  • Model Card

    A model card is a structured disclosure document published alongside an AI model that describes what the model is, how it was trained, what it is good at, where it fails, and what safety evaluations it passed. It gives engineers the information they need to decide whether a model fits their use case.

  • Model Context Protocol (MCP)

    Model Context Protocol (MCP) is an open standard that lets AI applications connect to external tools, data sources, and services through one common interface. Instead of writing a custom integration for every model and every tool, developers build an MCP server once and any MCP-compatible client, such as an AI assistant or coding agent, can use it.

  • Model Distillation

    Model distillation is the technique of training a smaller student model to reproduce the behavior of a larger teacher model, using the teacher's outputs as training signal. The student keeps most of the teacher's capability on the targeted tasks while running faster and cheaper, which makes it a standard tool for cutting inference cost. The canonical early result is DistilBERT, which Hugging Face reported in 2019 as 40% smaller and 60% faster than BERT while retaining 97% of its language understanding.

  • Model Drift

    Model drift is the degradation of an AI system's real-world performance over time while its code stays unchanged. It happens when the world shifts away from what the model learned (data and concept drift) or, in LLM-based systems, when a provider updates the underlying model and behavior changes beneath your prompts.

  • Multi-Agent System

    A multi-agent system is an architecture in which several AI agents with distinct roles work together on a problem too large or too varied for a single agent to handle well. Each agent runs in its own context with its own instructions and tools, and a coordination layer moves tasks and results between them.

  • Multimodal AI

    Multimodal AI refers to models that accept and reason over more than one type of input, such as text, images, audio, and video, within a single system. For software teams, it is the capability that lets an agent read a screenshot, interpret a diagram, or watch a screen recording instead of working from text alone. The capability arrived with force: GPT-4, a large multimodal model accepting image and text inputs, scored around the top 10% of test takers on a simulated bar exam in 2023.

O

  • Orchestrator Agent

    An orchestrator agent is the lead agent in a multi-agent system. It receives the overall goal, decomposes it into subtasks, delegates each one to worker agents with the context they need, evaluates what comes back, and assembles the pieces into a final result. It makes coordination decisions rather than doing the specialist work itself.

P

  • Prompt Caching

    Prompt caching is an inference optimization that stores the computed internal state of a repeated prompt prefix, so later API calls reusing that prefix skip reprocessing it. For agents that resend a large system prompt, tool definitions, and growing conversation history on every turn, it cuts input cost dramatically and reduces time to first token.

  • Prompt Chaining

    Prompt chaining is the technique of breaking a complex task into a sequence of separate model calls, where each prompt consumes the output of the previous one. Instead of asking for everything in a single prompt, each step does one focused job, trading extra latency and cost for higher reliability and easier debugging.

  • Prompt Engineering

    Prompt engineering is the practice of designing the text sent to an AI model, including instructions, examples, constraints, and output format, so the model reliably produces the result you want. It treats the prompt as an engineered artifact, something written deliberately, tested against cases, and versioned in source control, rather than a question typed casually into a chat box. The field is bigger than it looks from the outside: The Prompt Report, a 2024 survey by Schulhoff et al., catalogs 58 distinct text-based prompting techniques plus 40 more for other modalities.

  • Prompt Injection

    Prompt injection is an attack where malicious instructions are hidden inside content an AI system processes, such as a web page, email, or document, causing the model to follow the attacker's commands instead of the user's. It exploits the fact that language models cannot reliably separate trusted instructions from untrusted data. The OWASP Top 10 for LLM Applications 2025 ranks prompt injection as LLM01, the number-one risk on the list.

R

  • ReAct Agent

    A ReAct agent is an AI agent built on the Reasoning and Acting pattern: the model writes out a thought about what to do next, takes one action such as a tool call, observes the result, and repeats the cycle until the task is complete. Interleaving explicit reasoning with actions grounds each decision in real observations.

  • Reasoning Model

    A reasoning model is a large language model trained to work through a problem step by step before committing to an answer, spending extra inference compute on internal deliberation. That thinking phase, often hidden or summarized, raises accuracy on math, code, and multi-step planning tasks at the cost of higher latency and token spend. The gains are measurable: DeepSeek-R1 scored 79.8% pass@1 on the AIME 2024 math competition, slightly ahead of OpenAI's o1-1217.

  • Retrieval-Augmented Generation (RAG)

    Retrieval-augmented generation (RAG) is an architecture that searches a knowledge source for documents relevant to a query at request time and inserts them into the model's prompt, so the answer is grounded in retrieved evidence rather than only in what the model memorized during training. It gives models access to private, current, and verifiable information. The name comes from Lewis et al.'s 2020 NeurIPS paper at Meta AI, which set the state of the art on three open-domain question-answering tasks and produced language the authors measured as more specific, diverse, and factual than a parametric-only baseline.

  • RLHF

    RLHF, reinforcement learning from human feedback, is a training method that aligns a language model with human preferences. People rate or rank pairs of model outputs, a reward model learns to predict those judgments, and the language model is then optimized against that reward. It is the step that turned raw text predictors into usable assistants.

S

  • Semantic Search

    Semantic search is a retrieval method that matches queries to documents by meaning instead of shared keywords, typically by converting both into vector embeddings and ranking documents by how close their vectors sit to the query's. A search for "reset my password" can surface a document titled "account credential recovery" even though the two share no words. The approach reached mainstream scale in 2019, when Google deployed BERT in Search to interpret query meaning and said the model would improve understanding of one in 10 English searches in the U.S.

  • Shadow AI

    Shadow AI is the use of AI tools inside an organization without approval, oversight, or security review: employees pasting company data into consumer chatbots, wiring unvetted AI features into workflows, or running personal agent subscriptions on work tasks. It is the AI-era successor to shadow IT, with data exposure as the central risk.

  • Skill Atrophy

    Skill atrophy is the gradual decay of a developer's ability to write, read, and debug code unaided, caused by delegating that work to AI tools. The loss is invisible day to day because agent output keeps shipping on schedule, and it surfaces only when the AI fails, hits its limits, or produces something wrong that a human must untangle alone.

  • Slopsquatting

    Slopsquatting is a supply-chain attack where someone registers packages under names that AI coding tools hallucinate, so developers who install a model's invented dependency pull the attacker's code instead of an error. It works because models hallucinate the same plausible package names repeatedly, making those names predictable and worth squatting.

  • Small Language Model

    A small language model (SLM) is a compact language model, typically under about ten billion parameters, designed to run fast and cheap, often on a single GPU, a laptop, or a phone. It trades the peak capability of frontier models for low latency, low cost, data privacy, and the option to deploy on-device.

  • Spec-Driven Development

    Spec-driven development is the practice of writing a precise specification, with requirements and acceptance criteria, before an AI agent starts building. The spec becomes the durable source of truth that the agent implements against and reviewers verify against, replacing the scattered prompt history that ad hoc AI coding leaves behind.

  • Structured Outputs

    Structured outputs is a technique that forces a language model's response to conform to a developer-supplied schema, typically JSON Schema, so downstream code can parse it without defensive guessing. Instead of hoping the model formats its answer correctly, the API constrains generation so only schema-valid output is possible.

  • Subagent

    A subagent is a child agent that a lead agent spawns to handle one scoped task in a separate context window. It receives a brief, works independently with its own tools and instructions, and returns only its result. The parent keeps its own context clean while the subagent absorbs the exploratory noise. Anthropic's research system is built on this idea, spinning up 3 to 5 subagents in parallel, each in its own context window, to work a complex query.

  • Synthetic Data

    Synthetic data is artificially generated data, usually produced by a model rather than collected from real users or events, used to train, fine-tune, or test AI systems. Teams reach for it when real data is scarce, private, expensive to label, or missing the edge cases a model needs to learn from.

  • System Prompt

    A system prompt is the standing set of instructions given to a language model before any user input, defining its role, rules, tone, and constraints for the entire session. Unlike a user message, it persists across every turn and takes priority when instructions conflict, making it the primary control surface for model behavior.

T

  • Tokenization

    Tokenization is the process of splitting text into tokens, the small chunks a language model actually reads, using a fixed vocabulary learned from training data. Common words become single tokens while rare words break into fragments, and every model API meters cost and context limits in the tokens this process produces. As a reference point, Anthropic's tokenizer maps roughly 3.5 English characters to one token, with the ratio varying across languages.

  • Tool Calling

    Tool calling is a language model's ability to invoke external functions, such as running a command, querying an API, or reading a file, instead of only generating text. The model emits a structured request naming a tool and its arguments, the host application executes it, and the result feeds back into the model's next step.

  • Transformer Model

    A transformer model is the neural network architecture behind modern large language models. Its defining feature is attention: layers that let every token in a sequence directly weigh every other token when computing what comes next. Introduced by Google researchers in 2017, it replaced sequential architectures because it trains efficiently in parallel on GPUs. The original model announced itself with results: 28.4 BLEU on WMT 2014 English-to-German translation, beating the prior best, ensembles included, by more than 2 BLEU.

V

  • Vector Database

    A vector database is a datastore built to hold embeddings, the numeric vectors that represent the meaning of text, images, or code, and to answer similarity queries over them fast. Given a query vector, it returns the stored items closest in meaning, which makes it the retrieval backbone of RAG systems, semantic search, and agent memory.

  • Vector Embedding

    A vector embedding is a list of numbers that represents the meaning of a piece of text, code, or an image, produced by a machine learning model. Content with similar meaning gets numbers that sit close together in that numeric space, which lets software find related items by measuring distance instead of matching keywords.

  • Vibe Coding

    Vibe coding is building software by describing what you want in natural language and accepting the AI-generated code without reviewing or fully understanding it. You judge the result by whether it appears to work, trading rigor and comprehension for speed. The term comes from a 2025 observation by Andrej Karpathy that developers could now "forget that the code even exists."

Z

  • Zero-Shot Learning

    The prompt contents. Zero-shot provides instructions only; few-shot adds a handful of worked input-output demonstrations for the model to imitate. Zero-shot is cheaper per call and easier to maintain, since there are no examples to curate and keep consistent. Few-shot wins when format precision, unusual conventions, or edge-case handling matter, because demonstrations communicate those things more exactly than prose. Neither updates the model; both are inference-time techniques, and the practical workflow is to start zero-shot, measure, and add shots only where the measurements say to.

Let’s get in touch

Ready to build your product?

Book a consultation call to get a free project assessment
and scope estimation for your project.

Start your project