YFarmX logoYFarmX

Perceptron

Perceptron Mk1.5

Perceptron Mk1.5 reads text, images, video and audio for applications such as robots and smart glasses. It returns text, spatial annotations and tool calls through a hosted API.

Released 25 September 20261 min readLarge Language ModelsLast updated:

Editorial illustration of Perceptron Mk1.5

Key facts

25 Sep 2026
Released
36,864 tokens
Context
8,192 tokens
Maximum output
$0.15 / $1.50 per M
Input / output

Perceptron Mk1.5 reads text, images, video and audio for applications such as robots and smart glasses. It returns text, spatial annotations and tool calls through a hosted API.

Perceptron Mk1.5 is an AI model that reads text, images, video and audio and returns information that software can act on. Perceptron released it on 25 September 2026 for applications including drones, robotic quadrupeds, smart glasses and phones.

An embodied agent is software connected to a device that senses or acts in the physical world. Mk1.5 can identify where an object appears, track it through video and call a software tool. A robot’s wider control system still has to decide how to use that output.

The model returns text and annotations

Mk1.5 can return text, points, boxes, polygons, temporal annotations identifying sections of input video and object tracks. Its launch also describes web search and calls to other agents. These outputs help an application locate objects and connect observations to a task.

The current model documentation specifies a 36,864-token context window and 8,192-token maximum output. It accepts audio as an input; its documented outputs are text and structured information.

Hosted usage has separate input and output charges

Charge per million tokens Price
Input $0.15
Cached input $0.0375
Output $1.50

These are the rates in the model documentation checked on 8 October 2026. Developers should test their own camera views, motion and audio conditions: a useful result in one demonstration does not establish reliability across every physical setting.