mirror of
https://github.com/rohitg00/ai-engineering-from-scratch.git
synced 2026-10-02 01:54:39 +08:00
Expand glossary to 60+ terms, add 20 myth busters
New myths.md with 20 common AI misconceptions busted: "AI understands language", "more parameters = smarter", "you need a PhD in math", "temperature = creativity", "fine-tuning adds knowledge", etc. Glossary expanded from 25 to 60+ terms covering: Adam, autoregressive, CNN, contrastive learning, dropout, eigenvalue, encoder, epoch, GAN, hyperparameter, inference, inductive bias, KV cache, latent space, learning rate, loss function, mixed precision, MoE, NaN, normalization, overfitting, parameter, perplexity, quantization, ReLU, self-attention, SFT, softmax, temperature, transfer learning, VAE, weight decay, etc.
This commit is contained in:
+7
-6
@@ -1,10 +1,11 @@
|
||||
# Glossary
|
||||
|
||||
Myth vs. reality for AI terms. Organized alphabetically.
|
||||
Two resources to cut through the noise:
|
||||
|
||||
Each entry follows the pattern:
|
||||
- **What people say** — the common (often wrong) understanding
|
||||
- **What it actually means** — the real definition
|
||||
- **Why it's called that** — origin of the name (when interesting)
|
||||
- **[terms.md](terms.md)** - 60+ AI terms explained. What people say vs what it actually means.
|
||||
- **[myths.md](myths.md)** - 20 common AI misconceptions busted with what's actually going on.
|
||||
|
||||
See [terms.md](terms.md) for all entries.
|
||||
Each term entry follows the pattern:
|
||||
- **What people say** - the common (often wrong) understanding
|
||||
- **What it actually means** - the real definition
|
||||
- **Why it's called that** - origin of the name (when interesting)
|
||||
|
||||
@@ -0,0 +1,167 @@
|
||||
# AI Myths Busted
|
||||
|
||||
Common misconceptions about AI, ML, and deep learning. Each one explained with what's actually going on.
|
||||
|
||||
---
|
||||
|
||||
## "AI understands language"
|
||||
|
||||
**Reality:** LLMs predict the next token based on statistical patterns in training data. They have no understanding, no beliefs, no world model (that we can prove). They're very good at pattern matching across billions of examples. The output looks like understanding because the patterns are rich enough to cover most situations.
|
||||
|
||||
**Why it matters:** If you treat an LLM as a reasoning engine, you'll be surprised when it confidently says wrong things. If you treat it as a pattern matcher, you'll design better systems around it.
|
||||
|
||||
---
|
||||
|
||||
## "More parameters = smarter model"
|
||||
|
||||
**Reality:** A 7B parameter model trained on high-quality data with good techniques can outperform a 70B model trained on garbage. Chinchilla showed that most models were over-parameterized and under-trained. The quality and quantity of training data matters as much as model size. Phi-2 (2.7B) beat models 10x its size on many benchmarks.
|
||||
|
||||
**Why it matters:** Don't default to the biggest model. Match model size to your task and budget.
|
||||
|
||||
---
|
||||
|
||||
## "Neural networks are black boxes"
|
||||
|
||||
**Reality:** We have tools to understand what neural networks learn. Attention visualization shows what tokens the model focuses on. Probing classifiers reveal what information is stored in hidden representations. Mechanistic interpretability is finding actual circuits (induction heads, feature detectors). It's not complete transparency, but it's not a black box either.
|
||||
|
||||
**Why it matters:** You can debug neural networks. Gradient analysis, activation visualization, and attention maps are real tools covered in this course.
|
||||
|
||||
---
|
||||
|
||||
## "AI will replace programmers"
|
||||
|
||||
**Reality:** AI changed programming, it didn't replace it. AI writes boilerplate. Humans design systems, make architectural decisions, review correctness, and handle the cases AI gets wrong. The role shifted from "write every line" to "review, direct, and architect." The best engineers use AI as a tool, not fear it as a replacement.
|
||||
|
||||
**Why it matters:** You're learning AI engineering, which is programming + AI. Both skills together are more valuable than either alone.
|
||||
|
||||
---
|
||||
|
||||
## "You need a PhD in math to do AI"
|
||||
|
||||
**Reality:** You need high school math plus the specific topics in Phase 1 of this course. Linear algebra, calculus, probability, and optimization. You don't need proofs. You need intuition for what operations do and why they matter. If you can multiply matrices and take derivatives, you can build neural networks.
|
||||
|
||||
**Why it matters:** Phase 1 exists to give you exactly the math you need, nothing more.
|
||||
|
||||
---
|
||||
|
||||
## "GPT stands for General Purpose Technology"
|
||||
|
||||
**Reality:** GPT stands for Generative Pre-trained Transformer. Generative = it produces text. Pre-trained = trained once on a large corpus before being adapted. Transformer = the architecture from the 2017 "Attention Is All You Need" paper.
|
||||
|
||||
---
|
||||
|
||||
## "Temperature makes the AI more creative"
|
||||
|
||||
**Reality:** Temperature scales the logits before softmax. Higher temperature = flatter probability distribution = more random token selection. Lower temperature = sharper distribution = more deterministic. It's not creativity, it's randomness. A high-temperature model doesn't think harder, it just considers less likely tokens.
|
||||
|
||||
**Why it matters:** When your output is too repetitive, raise temperature. When it's too chaotic, lower it. It's a randomness knob, nothing more.
|
||||
|
||||
---
|
||||
|
||||
## "Fine-tuning teaches the model new knowledge"
|
||||
|
||||
**Reality:** Fine-tuning adjusts how the model uses existing knowledge, not what it knows. If information wasn't in the pre-training data, fine-tuning won't reliably add it. Fine-tuning is better for changing behavior (style, format, tone, task-specific patterns) than for adding facts. For new knowledge, use RAG.
|
||||
|
||||
**Why it matters:** If you need the model to know about your company's internal docs, use RAG. If you need it to respond in a specific format, fine-tune.
|
||||
|
||||
---
|
||||
|
||||
## "Bigger context window = better"
|
||||
|
||||
**Reality:** Models degrade on long contexts. The "lost in the middle" problem means models pay more attention to the beginning and end of long prompts and less to the middle. A 200K context window doesn't mean the model uses all 200K tokens equally well. Also, longer contexts cost more and are slower.
|
||||
|
||||
**Why it matters:** Don't dump everything into the context. Be selective. RAG with targeted retrieval beats stuffing the full document in.
|
||||
|
||||
---
|
||||
|
||||
## "AI agents are autonomous"
|
||||
|
||||
**Reality:** Current AI agents run in a loop: think, act, observe, repeat. They follow the pattern the harness defines. They don't have goals, plans, or self-awareness. They're reactive systems that use LLMs to decide what tool to call next. The "autonomy" comes from the loop, not from the AI.
|
||||
|
||||
**Why it matters:** When building agents, you're building the loop, the tools, and the guardrails. The LLM is just the decision-making component inside your system.
|
||||
|
||||
---
|
||||
|
||||
## "Transformers understand order because of positional encoding"
|
||||
|
||||
**Reality:** Transformers have no inherent sense of order. Self-attention treats input as a set, not a sequence. Positional encoding is a hack to inject order information by adding position-dependent vectors to the input. Different methods (sinusoidal, learned, RoPE, ALiBi) handle this differently. None of them truly give the model sequential understanding the way RNNs had it.
|
||||
|
||||
**Why it matters:** This is why positional encoding research is still active. It's a solved-enough problem for most uses, but it's fundamentally a workaround.
|
||||
|
||||
---
|
||||
|
||||
## "Pre-training is just reading the internet"
|
||||
|
||||
**Reality:** Pre-training is next-token prediction on a massive corpus. The model learns to predict what comes next given what came before. Through this simple objective, it learns grammar, facts, reasoning patterns, code structure, and more. But it also learns internet nonsense, biases, and incorrect information. The data curation, filtering, and deduplication matter enormously.
|
||||
|
||||
**Why it matters:** Garbage in, garbage out. The quality of pre-training data is one of the biggest differentiators between models.
|
||||
|
||||
---
|
||||
|
||||
## "RLHF aligns AI with human values"
|
||||
|
||||
**Reality:** RLHF aligns AI with the preferences of the specific humans who provided feedback. Those humans disagree with each other, have biases, and can't cover every situation. RLHF makes the model helpful and harmless in the ways the raters defined, not aligned with some universal human value system.
|
||||
|
||||
**Why it matters:** RLHF is a training technique, not a solution to alignment. It's one tool in a larger toolkit.
|
||||
|
||||
---
|
||||
|
||||
## "Embeddings capture meaning"
|
||||
|
||||
**Reality:** Embeddings capture statistical co-occurrence patterns. Words that appear in similar contexts get similar vectors. This correlates with meaning well enough to be useful, but it's not semantic understanding. "King - Man + Woman = Queen" works because of distributional patterns, not because the model understands monarchy or gender.
|
||||
|
||||
**Why it matters:** Embeddings are powerful for similarity search, clustering, and retrieval. But don't over-interpret what "similar" means.
|
||||
|
||||
---
|
||||
|
||||
## "Zero-shot means no training"
|
||||
|
||||
**Reality:** Zero-shot means no task-specific examples at inference time. The model was still trained on billions of tokens. It just hasn't seen examples of this specific task format. It generalizes from pre-training patterns. Few-shot means giving a few examples in the prompt. Neither means the model learned without training.
|
||||
|
||||
---
|
||||
|
||||
## "AI models learn like humans"
|
||||
|
||||
**Reality:** Humans learn from few examples, generalize across domains, and update beliefs continuously. Neural networks need millions of examples, generalize within their training distribution, and have fixed weights after training. The learning analogy is loose at best. Backpropagation is nothing like how biological neurons learn.
|
||||
|
||||
**Why it matters:** Don't anthropomorphize models. It leads to wrong expectations about what they can and can't do.
|
||||
|
||||
---
|
||||
|
||||
## "Scaling laws mean bigger is always better"
|
||||
|
||||
**Reality:** Scaling laws describe predictable relationships between compute, data, and model size. They show diminishing returns: doubling parameters doesn't double performance. They also assume you scale data proportionally. Many practical improvements come from better architectures, training techniques, and data quality, not just scale.
|
||||
|
||||
**Why it matters:** A 7B model with good engineering can solve your problem. Don't reach for 70B by default.
|
||||
|
||||
---
|
||||
|
||||
## "Open source AI is the same as open weights"
|
||||
|
||||
**Reality:** Most "open source" models are open weights. You get the model files but not the training data, training code, or data pipeline. True open source (like OLMo) releases everything: data, code, intermediate checkpoints, evaluation. Open weights is useful but not the same commitment as open source.
|
||||
|
||||
**Why it matters:** Know what you're getting. Open weights let you run and fine-tune. True open source lets you reproduce and understand.
|
||||
|
||||
---
|
||||
|
||||
## "Prompt engineering is not real engineering"
|
||||
|
||||
**Reality:** Prompt engineering is system design. You're designing the interface between human intent and model behavior. Good prompt engineering requires understanding tokenization, attention patterns, context window limits, and output parsing. It's closer to API design than to "talking nicely to the AI."
|
||||
|
||||
**Why it matters:** This course teaches prompt engineering as a real engineering discipline in Phase 11.
|
||||
|
||||
---
|
||||
|
||||
## "CNNs are outdated, everything is transformers now"
|
||||
|
||||
**Reality:** Vision Transformers (ViT) beat CNNs on many benchmarks, but CNNs are still used extensively. They're faster for inference, work well on mobile/edge, need less data, and have useful inductive biases (translation invariance, local patterns). Many production vision systems still use CNNs. The best architectures often combine both.
|
||||
|
||||
**Why it matters:** Learn both (Phases 4 and 7). Use what works for your constraints.
|
||||
|
||||
---
|
||||
|
||||
## "You need massive compute to train useful models"
|
||||
|
||||
**Reality:** You need massive compute to pre-train foundation models. But fine-tuning, LoRA, and transfer learning let you adapt models on a single GPU. Many useful AI applications don't require training at all, just good prompting and RAG. The "compute barrier" is for building foundation models, not for using them.
|
||||
|
||||
**Why it matters:** You can build real AI applications with a laptop. This course proves it.
|
||||
+181
-3
@@ -14,7 +14,15 @@
|
||||
|
||||
### Alignment
|
||||
- **What people say:** "Making AI safe"
|
||||
- **What it actually means:** The technical challenge of making an AI system's behavior match human intentions, values, and preferences — including edge cases the designer didn't anticipate
|
||||
- **What it actually means:** The technical challenge of making an AI system's behavior match human intentions, values, and preferences, including edge cases the designer didn't anticipate
|
||||
|
||||
### Autoregressive
|
||||
- **What people say:** "The AI generates one word at a time"
|
||||
- **What it actually means:** A model that predicts the next token conditioned on all previous tokens, then feeds that prediction back as input for the next step. GPT, LLaMA, and Claude are all autoregressive.
|
||||
|
||||
### Adam (Optimizer)
|
||||
- **What people say:** "The default optimizer"
|
||||
- **What it actually means:** Adaptive Moment Estimation. Combines momentum (first moment) with adaptive learning rates per parameter (second moment). Has bias correction for early steps. Works well across most tasks without much tuning.
|
||||
|
||||
## B
|
||||
|
||||
@@ -33,8 +41,28 @@
|
||||
- **What people say:** "Making the AI think step by step"
|
||||
- **What it actually means:** A prompting technique where you ask the model to show its reasoning steps, which improves accuracy on multi-step problems because each step conditions the next token generation
|
||||
|
||||
### CNN (Convolutional Neural Network)
|
||||
- **What people say:** "Image AI"
|
||||
- **What it actually means:** A neural network that uses convolution operations (sliding filters over the input) to detect local patterns. Stacking convolutions detects increasingly complex features: edges, textures, objects.
|
||||
|
||||
### CUDA
|
||||
- **What people say:** "GPU programming"
|
||||
- **What it actually means:** NVIDIA's parallel computing platform. Lets you run matrix operations on thousands of GPU cores simultaneously. PyTorch and TensorFlow use CUDA under the hood.
|
||||
|
||||
### Contrastive Learning
|
||||
- **What people say:** "Learning by comparison"
|
||||
- **What it actually means:** Training by pulling similar pairs closer and pushing dissimilar pairs apart in embedding space. CLIP uses this: matching image-text pairs vs non-matching ones.
|
||||
|
||||
## D
|
||||
|
||||
### Data Augmentation
|
||||
- **What people say:** "Making more training data"
|
||||
- **What it actually means:** Creating modified copies of existing data (rotate images, add noise, paraphrase text) to increase training set diversity without collecting new data. Reduces overfitting.
|
||||
|
||||
### Decoder
|
||||
- **What people say:** "The output part"
|
||||
- **What it actually means:** In transformers, a decoder uses causal (masked) self-attention so each position can only attend to earlier positions. GPT is decoder-only. BERT is encoder-only. T5 is encoder-decoder.
|
||||
|
||||
### Diffusion Model
|
||||
- **What people say:** "AI that generates images from noise"
|
||||
- **What it actually means:** A model trained to reverse a gradual noising process — it learns to predict and remove noise, and at generation time starts from pure noise and iteratively denoises
|
||||
@@ -43,15 +71,35 @@
|
||||
- **What people say:** "A simpler RLHF"
|
||||
- **What it actually means:** A training method that skips the reward model entirely — it directly optimizes the language model to prefer the better response in pairs of human preferences
|
||||
|
||||
### Dropout
|
||||
- **What people say:** "Randomly turning off neurons"
|
||||
- **What it actually means:** During training, randomly set a fraction of activations to zero. Forces the network to not rely on any single neuron. Turned off during inference. Simple but effective regularization.
|
||||
|
||||
## E
|
||||
|
||||
### Eigenvalue
|
||||
- **What people say:** "Some math thing for PCA"
|
||||
- **What it actually means:** For a matrix A, an eigenvalue lambda satisfies Av = lambda*v for some vector v. It tells you how much the matrix scales vectors in that direction. Large eigenvalues = directions of high variance in your data.
|
||||
|
||||
### Embedding
|
||||
- **What people say:** "Some AI magic that turns words into numbers"
|
||||
- **What it actually means:** A learned mapping from discrete items (words, images, users) to dense vectors in continuous space, where similar items end up close together
|
||||
- **Why it's called that:** The items are "embedded" in a geometric space where distance has meaning
|
||||
|
||||
### Encoder
|
||||
- **What people say:** "The input part"
|
||||
- **What it actually means:** In transformers, an encoder uses bidirectional self-attention so each position can attend to all positions. BERT is encoder-only. Good for understanding tasks (classification, NER) but not generation.
|
||||
|
||||
### Epoch
|
||||
- **What people say:** "One pass through the data"
|
||||
- **What it actually means:** Exactly that. One complete pass through every example in the training set. Multiple epochs = seeing the data multiple times. More epochs can improve learning but risks overfitting.
|
||||
|
||||
## F
|
||||
|
||||
### Feature
|
||||
- **What people say:** "A column in your data"
|
||||
- **What it actually means:** An individual measurable property of the data. In classical ML, you engineer features by hand. In deep learning, the network learns features automatically from raw data.
|
||||
|
||||
### Fine-tuning
|
||||
- **What people say:** "Training the AI on your data"
|
||||
- **What it actually means:** Starting with a pre-trained model's weights and continuing training on a smaller, task-specific dataset. Only updates existing weights, doesn't add new knowledge from scratch
|
||||
@@ -63,18 +111,54 @@
|
||||
- **What it actually means:** Generative Pre-trained Transformer — a specific architecture that predicts the next token using a decoder-only transformer trained on large text corpora
|
||||
- **Why it's called that:** Generative (produces text), Pre-trained (trained once on large data, then adapted), Transformer (the architecture)
|
||||
|
||||
### GAN (Generative Adversarial Network)
|
||||
- **What people say:** "Two AIs fighting each other"
|
||||
- **What it actually means:** A generator network tries to create realistic data while a discriminator network tries to tell real from fake. They train together: the generator gets better at fooling the discriminator, and the discriminator gets better at detecting fakes.
|
||||
|
||||
### Gradient
|
||||
- **What people say:** "The slope"
|
||||
- **What it actually means:** A vector of partial derivatives pointing in the direction of steepest increase. In ML, you go opposite to the gradient (gradient descent) to minimize the loss.
|
||||
|
||||
### Gradient Descent
|
||||
- **What people say:** "How AI improves"
|
||||
- **What it actually means:** An optimization algorithm that adjusts parameters in the direction that reduces the loss function most steeply, like walking downhill in a high-dimensional landscape
|
||||
|
||||
## H
|
||||
|
||||
### Hyperparameter
|
||||
- **What people say:** "Settings you tune"
|
||||
- **What it actually means:** Values set before training that control the training process itself: learning rate, batch size, number of layers, dropout rate. Unlike model parameters (weights), these aren't learned from data.
|
||||
|
||||
### Hallucination
|
||||
- **What people say:** "The AI is lying" or "making things up"
|
||||
- **What it actually means:** The model generates plausible-sounding text that isn't grounded in its training data or the given context — it's pattern-completing, not fact-retrieving
|
||||
|
||||
## I
|
||||
|
||||
### Inference
|
||||
- **What people say:** "Running the AI"
|
||||
- **What it actually means:** Using a trained model to make predictions on new data. No weight updates happen. This is what you do in production: send input, get output.
|
||||
|
||||
### Inductive Bias
|
||||
- **What people say:** Never heard of it
|
||||
- **What it actually means:** The assumptions built into a model's architecture. CNNs assume local patterns matter (convolution). RNNs assume order matters (sequential processing). Transformers assume everything might relate to everything (attention). The right bias helps the model learn faster from less data.
|
||||
|
||||
## K
|
||||
|
||||
### KV Cache
|
||||
- **What people say:** "Makes inference faster"
|
||||
- **What it actually means:** During autoregressive generation, caching the key and value matrices from previous tokens so you don't recompute them at each step. Trades memory for speed. Essential for fast LLM inference.
|
||||
|
||||
## L
|
||||
|
||||
### Latent Space
|
||||
- **What people say:** "The hidden representation"
|
||||
- **What it actually means:** A compressed, learned representation space where similar inputs map to nearby points. Autoencoders, VAEs, and diffusion models all work in latent space. It's lower-dimensional than the input but captures the important structure.
|
||||
|
||||
### Learning Rate
|
||||
- **What people say:** "How fast the AI learns"
|
||||
- **What it actually means:** A scalar that controls step size during gradient descent. Too high: overshoots the minimum and diverges. Too low: converges too slowly or gets stuck. The single most important hyperparameter.
|
||||
|
||||
### LLM (Large Language Model)
|
||||
- **What people say:** "AI" or "the brain"
|
||||
- **What it actually means:** A transformer-based neural network trained to predict the next token in a sequence, with billions of parameters, trained on internet-scale text data
|
||||
@@ -83,14 +167,54 @@
|
||||
- **What people say:** "Efficient fine-tuning"
|
||||
- **What it actually means:** Instead of updating all weights, insert small low-rank matrices alongside the original weights. Only these small matrices are trained, reducing memory by 10-100x
|
||||
|
||||
### Loss Function
|
||||
- **What people say:** "How wrong the AI is"
|
||||
- **What it actually means:** A function that measures the gap between predicted and actual output. Training minimizes this function. MSE for regression, cross-entropy for classification, contrastive loss for embeddings. The choice of loss function defines what "good" means to the model.
|
||||
|
||||
## M
|
||||
|
||||
### Mixed Precision
|
||||
- **What people say:** "Training trick for speed"
|
||||
- **What it actually means:** Using float16 for forward pass and most operations (faster, less memory) but keeping float32 for gradient accumulation and weight updates (more precise). Gets 2x speedup with negligible accuracy loss.
|
||||
|
||||
### MoE (Mixture of Experts)
|
||||
- **What people say:** "Only part of the model runs"
|
||||
- **What it actually means:** A model with many "expert" subnetworks where a routing mechanism sends each input to only a few experts. The full model is huge but each forward pass is cheap because most experts are skipped. Mixtral and GPT-4 use this.
|
||||
|
||||
### MCP (Model Context Protocol)
|
||||
- **What people say:** "A way for AI to use tools"
|
||||
- **What it actually means:** An open protocol (JSON-RPC over stdio/HTTP) that standardizes how AI applications connect to external data sources and tools, with typed schemas for tools, resources, and prompts
|
||||
|
||||
## N
|
||||
|
||||
### NaN (Not a Number)
|
||||
- **What people say:** "Training crashed"
|
||||
- **What it actually means:** A floating-point value indicating undefined results (0/0, inf-inf). In training, NaN loss usually means: learning rate too high, exploding gradients, log of zero, or division by zero. Always the first thing to check when training fails.
|
||||
|
||||
### Normalization
|
||||
- **What people say:** "Scaling the data"
|
||||
- **What it actually means:** Adjusting values to a standard range. Batch normalization normalizes across a batch. Layer normalization normalizes across features. Both stabilize training and allow higher learning rates.
|
||||
|
||||
## O
|
||||
|
||||
### Overfitting
|
||||
- **What people say:** "The model memorized the data"
|
||||
- **What it actually means:** The model performs well on training data but poorly on unseen data. It learned the noise, not the signal. Fix with: more data, regularization (dropout, weight decay), early stopping, data augmentation, simpler model.
|
||||
|
||||
### Optimizer
|
||||
- **What people say:** "The thing that updates weights"
|
||||
- **What it actually means:** An algorithm that uses gradients to update model parameters. SGD is the simplest. Adam is the most common. Each optimizer has different properties: convergence speed, memory usage, sensitivity to hyperparameters.
|
||||
|
||||
## P
|
||||
|
||||
### Parameter
|
||||
- **What people say:** "Model size"
|
||||
- **What it actually means:** A learnable value in the model, typically a weight or bias. "7B parameters" means 7 billion learnable numbers. Each float32 parameter takes 4 bytes, so 7B parameters = 28GB of memory just for the weights.
|
||||
|
||||
### Perplexity
|
||||
- **What people say:** "How confused the model is"
|
||||
- **What it actually means:** The exponential of the average cross-entropy loss. Lower is better. A perplexity of 10 means the model is as uncertain as if it were choosing uniformly among 10 tokens at each step.
|
||||
|
||||
### Prompt Engineering
|
||||
- **What people say:** "Talking to AI the right way"
|
||||
- **What it actually means:** Designing the input text to reliably produce desired outputs — including system prompts, few-shot examples, format instructions, and chain-of-thought triggers
|
||||
@@ -106,8 +230,28 @@
|
||||
- **What people say:** "How they make AI helpful"
|
||||
- **What it actually means:** A training pipeline: (1) collect human preferences on model outputs, (2) train a reward model on those preferences, (3) use PPO to optimize the LLM to produce higher-reward outputs
|
||||
|
||||
### Quantization
|
||||
- **What people say:** "Making the model smaller"
|
||||
- **What it actually means:** Reducing the precision of model weights from float32 (4 bytes) to int8 (1 byte) or int4 (0.5 bytes). Trades a small amount of accuracy for 4-8x less memory and faster inference. GPTQ, AWQ, and GGUF are common formats.
|
||||
|
||||
### ReLU
|
||||
- **What people say:** "Activation function"
|
||||
- **What it actually means:** Rectified Linear Unit: f(x) = max(0, x). The simplest non-linear activation. Fast to compute, doesn't saturate for positive values. Used everywhere because it works and is cheap. Variants: LeakyReLU, GELU, SiLU.
|
||||
|
||||
## S
|
||||
|
||||
### Self-Attention
|
||||
- **What people say:** "How the model decides what to focus on"
|
||||
- **What it actually means:** Each token computes query, key, and value vectors. Attention weight between two tokens = dot product of their query and key, scaled and softmaxed. Output = weighted sum of value vectors. Lets every token see every other token.
|
||||
|
||||
### SFT (Supervised Fine-Tuning)
|
||||
- **What people say:** "Teaching the model to follow instructions"
|
||||
- **What it actually means:** Fine-tuning a pre-trained model on (instruction, response) pairs. The model learns to generate the response given the instruction. This is what turns a base model into a chat model.
|
||||
|
||||
### Softmax
|
||||
- **What people say:** "Turns numbers into probabilities"
|
||||
- **What it actually means:** softmax(x_i) = exp(x_i) / sum(exp(x_j)). Transforms a vector of arbitrary real numbers into a probability distribution (all positive, sums to 1). Used in classification heads, attention weights, and anywhere you need probabilities.
|
||||
|
||||
### Swarm
|
||||
- **What people say:** "A bunch of AI agents working together like bees"
|
||||
- **What it actually means:** Multiple agents sharing state and coordinating through message passing, with emergent behavior arising from simple individual rules rather than central control
|
||||
@@ -118,13 +262,47 @@
|
||||
- **What people say:** "A word"
|
||||
- **What it actually means:** A subword unit (typically 3-4 characters in English) produced by a tokenizer like BPE. "unbelievable" might be 3 tokens: "un" + "believ" + "able"
|
||||
|
||||
### Temperature
|
||||
- **What people say:** "Creativity setting"
|
||||
- **What it actually means:** A scalar that divides logits before softmax. Temperature=1 is default. Higher = flatter distribution = more random outputs. Lower = sharper distribution = more deterministic. Temperature=0 is argmax (always pick the most likely token).
|
||||
|
||||
### Transfer Learning
|
||||
- **What people say:** "Using a pre-trained model"
|
||||
- **What it actually means:** Taking a model trained on one task and adapting it to a different task. The early layers learn general features (edges, syntax patterns) that transfer. Only the later layers need task-specific training. This is why you can fine-tune BERT for any NLP task.
|
||||
|
||||
### Transformer
|
||||
- **What people say:** "The architecture behind modern AI"
|
||||
- **What it actually means:** A neural network architecture that processes sequences using self-attention (letting every position attend to every other position) instead of recurrence, enabling massive parallelization
|
||||
- **Why it's called that:** It transforms input representations into output representations through attention layers — and the name sounded cool
|
||||
- **Why it's called that:** It transforms input representations into output representations through attention layers
|
||||
|
||||
## U
|
||||
|
||||
### Underfitting
|
||||
- **What people say:** "The model isn't learning"
|
||||
- **What it actually means:** The model is too simple to capture the patterns in the data. Training loss stays high. Fix with: more parameters, more layers, longer training, lower regularization, better features.
|
||||
|
||||
## V
|
||||
|
||||
### VAE (Variational Autoencoder)
|
||||
- **What people say:** "A generative model"
|
||||
- **What it actually means:** An autoencoder that learns a smooth latent space by forcing the encoder output to follow a Gaussian distribution. You can sample from this distribution and decode to generate new data. The reparameterization trick makes it trainable via backpropagation.
|
||||
|
||||
### Vector Database
|
||||
- **What people say:** "A special database for AI"
|
||||
- **What it actually means:** A database optimized for storing vectors (dense arrays of floats) and performing fast approximate nearest-neighbor search — the core operation in similarity search, RAG, and recommendation systems
|
||||
- **What it actually means:** A database optimized for storing vectors (dense arrays of floats) and performing fast approximate nearest-neighbor search. The core operation in similarity search, RAG, and recommendation systems.
|
||||
|
||||
## W
|
||||
|
||||
### Weight
|
||||
- **What people say:** "What the model learned"
|
||||
- **What it actually means:** A single number in a model's parameter matrix. A linear layer with input size 768 and output size 3072 has 768*3072 = 2,359,296 weights. Training adjusts each weight to minimize the loss function.
|
||||
|
||||
### Weight Decay
|
||||
- **What people say:** "Regularization"
|
||||
- **What it actually means:** Adding a penalty proportional to the magnitude of weights to the loss function. Equivalent to L2 regularization. Prevents weights from growing too large. Typical value: 0.01-0.1.
|
||||
|
||||
## Z
|
||||
|
||||
### Zero-Shot
|
||||
- **What people say:** "No training needed"
|
||||
- **What it actually means:** Using a model on a task it wasn't explicitly trained for, with no task-specific examples in the prompt. The model generalizes from pre-training. Works because large models have seen enough variety to handle new task formats.
|
||||
|
||||
Reference in New Issue
Block a user