Latency
AILatency: The delay between sending a request and receiving the result; in language models it covers the wait for the first token and for the full response.
Used in these stories
- ChatGPT Images 2.5 adds sketch-led creation and more precise edits
- DeepSeek plans at least 160,000 Huawei Ascend chips in Inner Mongolia
- Oracle is installing Quantinuum's Helios quantum computer inside its AI cloud
- Microsoft rebuilds its AI so that any model can be swapped out
- OpenAI launches Presence, and it already answers 75% of OpenAI's own support calls
- Nvidia Ising decoder cuts colour code error rates by 347 times
- GPT-Live Voice brings full duplex conversation to ChatGPT
- Robostral Navigate marks Mistral's move into embodied AI
- IBM to build India quantum computer in Amaravati by September
- OpenAI Releases GPT-5.4 Mini and Nano, Its Smallest and Fastest GPT-5.4 Variants