--- okf_version: "0.2" --- # jimmylearning.com — Knowledge Base > How large language models really work — tokens, attention, training, sampling and scale — with hands-on interactive tools and the original research papers. A portable OKF/LLM-wiki bundle of the site's knowledge. Read this index, then open the relevant concept files. No embeddings or vector database required. ## Concepts * [AI, at a glance](concepts/newsroom.md) - One view of what's happening right now — headlines, fresh research and what people are arguing about — pulled from public sources on the server and refreshed on its own, so… * [Ask anything about AI](concepts/ask.md) - The field moves faster than any one page can. * [Chat with this site](concepts/chat.md) - A chatbot that answers from this site's own knowledge — not the open web. * [The model landscape](concepts/models.md) - Who builds the frontier, in one view. * [Fresh from Hugging Face](concepts/hfmodels.md) - The map above is the who; this is what shipped lately. * [How fast is this actually moving?](concepts/pace.md) - Three numbers tell the story of the last few years better than any headline: how much a model can read at once, how much compute went into training it, and how little it now… * [It never sees letters. It sees tokens.](concepts/tokens.md) - Before anything else happens, your text is chopped into pieces from a fixed vocabulary. * [Every token becomes a direction in space](concepts/embeddings.md) - Token 3,290 means nothing. * [Attention is a weighted lookup, and that is the whole trick](concepts/attention.md) - Each position asks a question, every earlier position answers, and the answers get averaged in proportion to how well they match. * [One forward pass, start to finish](concepts/forward-pass.md) - Here is the entire journey from your text to a single next token. * [Three passes, and only one of them is expensive](concepts/training.md) - A base model and a chat assistant are the same network at different points in the same process. * [Why the same question gives different answers](concepts/sampling.md) - The network's output is not a word. * [Scale is a curve someone measured](concepts/scale.md) - The decision to build very large models was not a hunch. * [The failure modes are structural, not bugs](concepts/limits.md) - Each of these follows directly from something described above. * [Six things that are commonly said](concepts/myths.md) - Each of these is a reasonable inference from watching a model behave. * [The questions people actually ask](concepts/faq.md) - Short answers to the things that bring most people here. * [The vocabulary, in one place](concepts/glossary.md) - The vocabulary, in one place * [Read the primary sources](concepts/sources.md) - Every claim above traces to one of these. * [Try a model, right here](concepts/playground.md) - Reading about models only goes so far — use one. * [Watch a real model think](concepts/lab.md) - The sampler in Part 06 used stand-in numbers. * [Explore Hugging Face, live](concepts/agent.md) - Ask in plain language — "what are the latest speech-to-text models?", then "show me one of them". * [Read — explainers, opened right here](concepts/library.md) - Hand-picked companion pieces that open in a reader on this page, so you can follow a thread without losing your place. * [Eight years, briefly](concepts/timeline.md) - Everything in this primer sits on a short, fast history. ## Glossary * [token](glossary/token.md) - The unit a model reads and writes. * [embedding](glossary/embedding.md) - The vector a token ID is converted into. * [logit](glossary/logit.md) - A raw, unnormalised score for one vocabulary entry, before softmax turns it into a probability. * [softmax](glossary/softmax.md) - Converts a list of scores into positive numbers that sum to 1. * [attention head](glossary/attention-head.md) - One query/key/value lookup. * [context window](glossary/context-window.md) - The maximum number of tokens the model can attend to at once. * [temperature](glossary/temperature.md) - Divides logits before softmax. * [top-p / nucleus](glossary/top-p-nucleus.md) - Keeps only the smallest set of candidates whose probability sums to p; discards the tail. * [RLHF](glossary/rlhf.md) - Reinforcement learning from human feedback. * [RAG](glossary/rag.md) - Retrieval-augmented generation. * [KV cache](glossary/kv-cache.md) - Stored keys and values for tokens already processed, so generating each new token does not recompute the whole sequence. * [mixture of experts](glossary/mixture-of-experts.md) - Routes each token through a small subset of many parallel sub-networks. ## FAQ * [How does a large language model work?](faq/how-does-a-large-language-model-work.md) - It converts your text into tokens, turns each token into a vector, passes those vectors through dozens of transformer layers that let every position read from earlier… * [What is a token in a language model?](faq/what-is-a-token-in-a-language-model.md) - A token is a chunk of text from a fixed vocabulary — usually a common word, a fragment of a rarer word, or punctuation. * [Why do language models hallucinate?](faq/why-do-language-models-hallucinate.md) - Because the training objective rewards plausible text, not true text. * [What does temperature do in an LLM?](faq/what-does-temperature-do-in-an-llm.md) - Temperature divides the model's raw scores before they become probabilities. * [What is attention in a transformer?](faq/what-is-attention-in-a-transformer.md) - A weighted lookup. * [Do language models remember previous conversations?](faq/do-language-models-remember-previous-conversations.md) - No. * [Should I use RAG or fine-tuning?](faq/should-i-use-rag-or-fine-tuning.md) - Use retrieval for knowledge — anything that changes, needs citing, or must be current. * [How much data does it take to train a large language model?](faq/how-much-data-does-it-take-to-train-a-large-language-model.md) - The 2022 Chinchilla result put the compute-optimal ratio at roughly 20 training tokens per parameter — a 70B model on about 1.4 trillion tokens. ## Reference * [Primary sources](reference/sources.md) - 12 foundational papers