The model is memoryless. The context has a length limit.
token — the unit of text, roughly a word, you pay per token
prompt — the part of the model input you control, context — the entire model input
GPT-6 Astra, Fable 5.1, Gemini 3.1 Pro, open weights: Qwen 3.8, DeepSeek-V4, Kimi 3, GLM-5.2, Gemma 4