Prefix caching
AIPrefix caching: A serving optimisation that keeps the cached computation for a shared opening sequence, so many requests beginning the same way skip repeated prefill work.
Prefix caching: A serving optimisation that keeps the cached computation for a shared opening sequence, so many requests beginning the same way skip repeated prefill work.