The path to 1M context window
How LLMs went from GPT-3's 2,048 tokens to 1M+ — positional encoding, attention rebuilt around sparse and linear layers, training data, optimizers, long-context RL, and the evals that separate advertised from effective context. With figures.