On May 27, 2026, Sesame — the voice AI startup founded by Oculus’ founders and early team members — put an iOS preview on the App Store in 39 countries, free during the initial rollout. It lands exactly one year after the company’s web Research Preview, which drew more than a million people to its voice agents Maya and Miles within the first few weeks, according to investor Sequoia. The app adds two new characters, Simone and Charlie, for a four-agent roster, each with its own voice, personality, point of view, and memory.
Unlike the ChatGPT-style “input box plus conversation history” form factor, Sesame was built voice-first from day one: the whole experience is designed around the rhythm and timing of natural speech, with visuals as an optional layer. The company insists on calling these agents, not chatbots.
Four Agents, Each With Its Own Memory
Sesame’s design bet is character-first. Maya, Miles, Simone, and Charlie were each deliberately crafted with a distinct voice, personality, and point of view — and their memories do not cross. What you told Miles stays with Miles. Memory carries over between voice and text, so switching to typing does not mean re-explaining your context. An Incognito Mode lets agents see prior context but store nothing new; the conversation is ephemeral. The design deliberately imitates real social relationships: you talk about different things with different people, and switching agents means switching context, not pouring everything into one faceless assistant. Whether that actually drives retention is the part of the bet still waiting to be tested.
Searching While Speaking: Parallel Retrieval and Mid-Sentence Pivots
The technically interesting part is latency. Sesame built fast search and retrieval systems and runs multiple parallel searches while the agent is still speaking, weaving fresh results into the response mid-stream — even pivoting mid-sentence. The company’s blog says this lets agents run “slower and smarter agentic loops without awkward interruptions.” As TechCrunch quotes the team: “There’s an inherent tension between replying quickly and taking the time to compose thoughtful responses,” and “a slower response is usually more correct, but it can also feel unnatural if it takes too long.”
The interface stays voice-first: search cards with image results, Notes for takeaways, Deep Dives for in-depth results, and a texting mode for when speaking aloud isn’t practical. Architecturally, this hides thinking time inside the speech stream — instead of dead air while retrieval finishes, the agent starts talking and corrects course as facts arrive. If it works, it rewrites both the perceived-latency budget and the retrieval cost of every conversation.
Free, 39 Countries, and a Million-User Head Start
Commercially, the app is free “for the time being,” with a short waitlist possible at sign-up and an Android preview on the way. The war chest is substantial: Sesame announced a $250 million Series B led by Sequoia in October 2025 alongside the beta opening, and Sequoia says the Research Preview passed a million users in its first weeks. By consumer voice-AI standards that is a strong opening, and the benchmark Sesame sets for itself is explicitly the traditional text-chatbot experience that ChatGPT popularized.
Launching across 39 countries at once also says something: voice products are far less forgiving of accents and localization gaps than text, so a first wave this wide implies real confidence in multilingual performance — or at least a willingness to place the bet on voice early.
From Thinking With You to Doing For You — and 2027 Eyewear
Sesame calls the iOS app a first step toward intelligent eyewear planned for 2027. For a team that built its career in VR, the route makes sense: on a glasses form factor, voice is the only natural input. The company also signals a coming shift — agents that can “do for you,” not just “think with you,” which is exactly why it insists on the word agent.
For product teams, two signals are worth tracking. One: whether character-based, per-agent memory design genuinely holds users, or collapses into a novelty. Two: whether “keep searching while speaking” becomes the standard architecture for voice products, because it changes the latency users will tolerate and the cost profile of every session. Monetization remains open — subscriptions, enterprise licensing, or hardware bundling will have to reach the table eventually, and the 2027 eyewear date will reveal who in this voice-agent wave is actually serious.
Sources
- Sesame, the conversational AI startup from Oculus founders, launches its iOS app — TechCrunch
- Voice your curiosity — Sesame
AI-assisted summary compiled from the sources above, reviewed by a human before publishing.
