OpenAI’s move into consumer hardware is making serious waves this week, and honestly, I couldn’t be more thrilled. According to recent reports, OpenAI has partnered with design legend Jony Ive and his LoveFrom studio to build a screenless, portable smart speaker — pitched as a “humanlike AI companion” that brings ChatGPT out of the browser and into the physical world. The goal: less time on our phones, more time talking.
What fascinates me most isn’t the rumored mechanical moving parts or the $200–$300 price point. It’s the fundamental shift in how the device intends to interact with users. Instead of the traditional command-and-response model built around a standard wake word, this device is reportedly aiming for an “observe-and-assist” approach — powered by the GPT-Live model, using cameras and microphones to read its surroundings and proactively engage.
It’s an exciting vision, and one that aligns with what we’ve been saying at Sensory for years about the future of natural human-machine interaction. But after 30 years putting speech recognition into deeply embedded hardware, I know firsthand that bringing ambient intelligence to a portable device runs into two massive engineering hurdles: power consumption and privacy.
A truly portable companion runs on a rechargeable battery. An “ambient” and proactive one has to constantly monitor video and audio to know when to engage. Running a camera, an open microphone, and a main processor continuously — streaming all of it to the cloud — is a recipe for a dead battery well before lunch.
The alternative isn’t much better on its own: push everything on-device, and you need bigger models and beefier processors, which drives up cost and reintroduces the same power and heat problems from a different direction.
The real fix is neither “bigger battery” nor “bigger chip.” It’s being selective about what actually needs to run continuously, and what only needs to wake up when it matters.
An always-listening, always-watching device that anticipates your needs using your emails, schedule, and habits is a privacy minefield. If all of that gets processed in the cloud, consumer pushback will be immediate and loud.
There’s also an identity problem: without a wake word, how does the device know who it’s talking to? If it’s pulling from your personal ChatGPT account, it needs to be confident it’s you talking — not your spouse, a guest, or a voice on the TV.
We don’t think the fix is a bigger battery or a thicker privacy policy. It’s smart, on-device processing paired with carefully screened cloud requests — which is exactly where Sensory’s technology fits, and how we help the world’s biggest brands build low-power, private voice interfaces.
Ultra-low-power wake word and intent detection. Instead of streaming ambient noise to the cloud around the clock, our wake word technology can run on just microwatts of power, listening for specific environmental cues, user presence, or custom triggers. It acts as an efficient gatekeeper, waking the power-hungry processor and cloud connection only when a genuine interaction is intended.
On-device speaker verification. To solve the identity problem, Sensory offers biometric speaker identification entirely on the edge. The device knows who’s speaking and unlocks personal context locally — voice biometrics never leave the device, keeping identity and personal data protected. Sensory supports this through wake-word-based (text-dependent) or free-speech (text-independent) verification, as well as on-device facial recognition.
To the teams at OpenAI and LoveFrom: we love the vision for a true AI companion — it’s the next great leap for our industry. Sensory has spent decades solving exactly the embedded challenges you’re facing right now, putting voice technology into billions of shipping consumer electronics. We’d love to help make this device not just intelligent, but power-efficient, responsive, and uncompromisingly private. Let’s talk!