Technology
Plain-language glossary of the technologies behind AI models and tools: transformers, mixture-of-experts, reasoning, RAG, MCP, quantization and more.
- Dense transformer — The standard attention-based architecture behind most LLMs. “Dense” means the whole network runs for each token — simple to train and serve, but compute grows with total size.
- Mixture of experts (MoE) — A router sends each token to a small subset of “expert” layers, so a model can hold trillions of parameters while only activating a fraction per token. Most frontier open-weight models are sparse MoEs — compare total vs active parameters.
- Hybrid state-space (Mamba) — State-space model (SSM) layers such as Mamba process sequences in linear time with constant memory. Hybrids interleave them with a few attention layers to keep recall while cutting long-context cost and boosting throughput.
- Sparse & linear attention — Variants that avoid full quadratic attention — sparse attention picks which tokens to attend to, linear attention compresses history into a fixed-size state. Used to make million-token contexts affordable.
- Diffusion language model — Instead of writing one token at a time, a diffusion LM starts from noise/masks and iteratively denoises the whole answer in parallel — trading some quality for very high output speed.
- Diffusion & flow matching — The dominant approach for media generation: a model learns to reverse a noising process (diffusion) or follow a learned velocity field (flow matching), usually with a transformer backbone (DiT).
- JEPA — Joint-Embedding Predictive Architecture (Yann LeCun / Meta). Rather than generating every pixel, a JEPA predicts the abstract embedding of missing or future parts of the input — learning what matters about how the world behaves. V-JEPA models are used as world models for understanding, prediction and robot planning.
- System One decisions — A model class (introduced by TypeSafe AI with Jev) built for the fast, intuitive “System 1” decisions inside software: unstructured state goes in, a schema-defined choice with calibrated probabilities comes out. Uses a parallel sampler and Reinforcement Learning for Calibrated Decisions (RLCD); outputs cannot fall outside the schema.
- Vision-language-action — VLA models extend vision-language models with an action output head, so one network can perceive a scene, follow a natural-language instruction and emit motor commands.
- Native multimodality — Rather than bolting separate encoders onto a text model, native (“omni”) models are trained on mixed modalities from the start, so they can reason across — and often generate — images, speech and video.
- Reasoning (test-time compute) — Models trained with reinforcement learning to produce a hidden chain of thought before the answer. Spending more tokens “thinking” (the effort setting) buys accuracy on hard problems at the cost of latency and price.
- Structured outputs — At each decoding step, tokens that would break the schema are masked out, so the response is guaranteed to parse. The basis of reliable function calling and typed extraction.
- Fine-tuning (LoRA) — Parameter-efficient methods such as LoRA train small adapter matrices instead of all weights, making it practical to specialise open models on a single GPU.
- Quantization — Storing weights at lower precision (FP8, INT4, GGUF formats) cuts memory and speeds up inference with modest quality loss — what makes local and on-device LLMs practical.
- Speculative decoding — A fast draft model guesses several tokens ahead and the target model checks them in one pass, giving identical output several times faster. Widely used by inference providers.
- Custom AI silicon — Non-GPU accelerators designed around LLM inference — such as Groq’s LPU and Cerebras’ wafer-scale engine — that keep weights in on-chip memory to reach very high tokens per second.
- Retrieval-augmented generation — Retrieve relevant passages (via vector, keyword or web search) and put them in the prompt, so the model answers from your data with citations instead of from memory.
- Vector search — Approximate nearest-neighbour indexes (HNSW, IVF, DiskANN) that search millions of embeddings in milliseconds — the retrieval layer of most RAG systems.
- Model Context Protocol — MCP defines how AI applications discover and call external tools, resources and prompts through MCP servers — one integration works across many assistants and agent frameworks.
- Agent orchestration — Frameworks that run a model in a loop with tools, memory and state (ReAct-style), and coordinate several specialised agents with handoffs, graphs or supervisors.
- Computer & browser use — Models operate real software through screenshots and mouse/keyboard actions or a controlled browser, automating tasks that have no API.
- Agent memory — Stores facts, preferences and past interactions outside the context window (often as a knowledge graph or vector store) so agents can recall them later.
- Evals & LLM-as-judge — Datasets, scorers and strong-model judges that measure quality, regressions and safety of LLM apps before and after release.
- Guardrails — Classifiers, validators and rules that block prompt injection, jailbreaks, PII leaks and unsafe output around an LLM call.
- Real-time voice — Streaming speech recognition, LLM and speech synthesis (or a single speech-to-speech model) wired together with turn detection and interruption handling for natural voice agents.
- Decision modeling (DMN) — Decisions expressed as rules, decision tables and decision graphs (often in the DMN standard) and served as decision APIs — deterministic and explainable, and increasingly combined with ML or LLM steps.