Flagship LLM, strong agent & coding capability, 500K context, text‑image input, self‑verification for long tasks