Glossary/InferentiaInferentiaAIInferentia: Amazon's in-house inference chip, aimed at cutting the cost per token of serving models on AWS.Used in these storiesOpenAI Jalapeno chip signals a full stack ambition25 June 2026Related termsInference-time scalingInfiniBandInference serverInference engineInferenceInductive biasBrowse the full glossary →