Latency
AILatency: The delay between sending a request and receiving the result; in language models it covers the wait for the first token and for the full response.
Used in these stories
- ChatGPT Images 2.5 adds sketch-led creation and more precise edits
- DeepSeek plans at least 160,000 Huawei Ascend chips in Inner Mongolia
- Oracle is installing Quantinuum's Helios quantum computer inside its AI cloud
- Microsoft rebuilds its AI so that any model can be swapped out
- OpenAI launches Presence, and it already answers 75% of OpenAI's own support calls
- Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber: Google's security model is the one to watch
- Nvidia Ising Decoder Cuts Colour Code Error Rates By 347 Times
- GPT-Live Voice Brings Full Duplex Conversation To ChatGPT
- Robostral Navigate Marks Mistral's Move Into Embodied AI
- IBM To Build India Quantum Computer In Amaravati By September