What happened
- On September 25, Perceptron released Mk1.5, a model for controlling agents that operate on physical objects, according to the announcement on its own blog. It takes text, images, video and audio, and returns text, points, boxes, polygons, crops and object trajectories.
- The company trained it to output the trajectory as timestamped geometry, rather than a standalone detection per frame. With that, it says it leads three of the four video object segmentation benchmarks it measured: Molmo2-Track, Ref-DAVIS17 and ReasonVOS. On MeViS, MolmoPoint-8B came out ahead.
- On first-person video, the announcement claims Mk1.5 locates hands 50% better than the strongest Gemini model they measured.
- It’s available now through the company’s platform with a 32,000-token multimodal context, priced at $0.15 per million input tokens and $1.50 per million output tokens. The company itself says it has deployed it on drones, robot dogs, glasses and phones.
Why it matters
- The use case the company names first isn’t industrial robotics: it’s in-store inventory and tracking players in a sports broadcast. For a retailer or a brand that sponsors sports in Chile, that lowers the cost of measuring what moved off the shelf or how long a logo was on camera.
- The concrete operational promise is eliminating the re-identification pipeline that currently sits behind any detector. Fewer in-house pieces to maintain means less development, and also less control over the criteria by which the system decides two images are the same object.
- The figures are the company’s own, measured by the company itself. There’s no independent verification, and the comparison chooses which models go into the table.
The number
$0.15 per million input tokens, with a 32,000-token multimodal context.
Context
The control layer for physical bodies is filling up fast: NVIDIA has already turned agents into reusable skills inside the robot, and Alphabet released its robotics foundation under an Apache license. What changes here is the price of perception, not of movement.
What’s next
- No timelines announced. The model is already published, and the company hasn’t committed to dates for subsequent versions or for deployment on customers’ own infrastructure.
Bottom line
A model that tracks objects for $0.15 per million makes cheap a question that wasn’t asked before because of cost: what’s happening in front of the camera that’s already installed.
Sources
- Introducing Perceptron Mk1.5, Perceptron blog, September 25, 2026.
Edited by Rodrigo Cornejo. How we select and verify: who writes these notes.
%2022.41.55.s5Pwg7YZ_1SFzdB.webp)

