Multimodal frontier models: Gemini 1.5 Pro & Flash with 2M token context windows.
LPU Inference Engine delivering 500+ tokens/second for open-weight LLMs.