Spatial computing needed something to hold
A components company, not a platform holder, has shipped a 299 dollar controller for Vision Pro. That gap is the interesting part, and it matches what we keep finding on site.
On 29 July DFRobot put the seeMote Cube on sale at 299 dollars. It is a handheld input device for Apple Vision Pro development: six degrees of freedom of motion tracking, six programmable physical buttons, and haptic feedback, in a 56mm block that talks to the headset over Bluetooth. The company points it at spatial design review, digital twin and model inspection, and training simulation (PR Newswire, 2026; LinuxGizmos, 2026).
The device is not the story. The story is who shipped it. This is a components company filling a gap the platform holder left open, three years into a product whose entire input pitch was that you would not need a controller.
Hand tracking demos well and works badly
We build interactive installations for public spaces, and we have run enough of them to be blunt about this. Camera-based hand tracking is excellent in a quiet room with even light and one motivated user. It degrades in exactly the conditions a real deployment has.
- 01Venue lighting is either a spotlight or a dim corner, and both wreck the tracking volume.
- 02Arms get tired. Anything that takes more than about ninety seconds of held-up hands loses people.
- 03People are carrying things. A coffee, a bag, a child. One hand is already gone.
- 04There is no tactile confirmation, so users repeat the gesture, and the system reads two inputs.
- 05A stranger has no idea what gestures exist, and nothing on screen can teach them fast enough.
An object is an interface and a permission model
Hand a person a physical thing and three problems solve themselves at once. They know they are now the one interacting, because they are holding the controller. They know roughly what they can do, because they can feel how many buttons there are. And they get a click back when something registers, which is the cheapest confidence signal in interaction design.
That third one matters more than it sounds. Most of the confusion we see on installation floors is not people failing to perform a gesture. It is people performing it correctly and not believing it worked.
The controller is not a fallback for when tracking fails. It is the part of the system that tells someone they are allowed to touch it.
What we would do with it
For a headset-based client build today, we would spec a held device from the start rather than adding one after the first user test goes badly. Coarse navigation and object manipulation on the motion tracking, every commit and cancel on a physical button, haptics on state changes only, and never on hover. The headset handles what you are looking at. The object in your hand handles what you decide.
The wider read is about where XR is actually maturing. The headline hardware has been iterating on displays and passthrough for years while input stayed a philosophical position. A 299 dollar accessory from outside the platform is the sound of that position getting quietly renegotiated by the people who have to ship working things.
If you are scoping something spatial and the plan says hands only, we would want to see it in the actual room before agreeing.
