You notice you're out of rice and chickpeas mid-cooking — and only order rice, forgetting chickpeas. The job isn't ordering; it's keeping track. So I built a voice-first agent: say what you need, and it finds your usual brands and builds the cart for you.
← Back to portfolioSay "add a litre of milk and some arhar dal" — in English, Hindi, or a mix. It transcribes Indian speech and brand names accurately.
An LLM agent plans the order: checks your saved address, your usual brands, and your preferences, then searches and picks the right products.
It assembles the basket and reads it back in a natural voice — you review and confirm before anything is bought.
of professionals said grocery, food & beverage is the number-one category they'd trust an AI agent to complete entirely on their behalf — ahead of travel, electronics, and every other category.
said their biggest worry is that the agent will get it wrong (and 20% don't want to give up control) — exactly why this agent builds the cart but leaves the final confirm to you.
Data based on survey of 1000+ US professionals.
A 0→1 build I did solo — both the working product and the commercial thinking behind it:
Voice-to-cart agent: speech-to-text → an LLM tool-calling agent (Model Context Protocol) → streaming voice, with a personalization profile and multilingual (Hindi/English) support.
Per-order and per-user unit-economics model, a customer-validation survey, and a competitive teardown — to test whether it's a real, monetizable product.
Spec'd a low-cost kitchen device (ESP32 + mic + speaker) that turns the phone into the "brain" — the route from software demo to physical product.
Same agent architecture family (Claude + MCP) used by conversational-checkout products in agentic commerce today.
Voice input, an MCP-based tool-calling agent, and a live commerce API doing the checkout — that's the shape most conversational commerce agents take today: a reasoning layer that plans the order, and a checkout layer that executes it. The reasoning layer is largely solved — Claude- and GPT-class models with MCP-style tool calling handle intent and product matching well. The harder, less-solved piece is the checkout layer itself, specifically agent-native payment authorization: letting an agent transact within a user's pre-approved limits without a human clicking confirm every time. This project keeps that human confirm step deliberately, since spending-limit infrastructure at the payment-rail level is still early across most markets.
Indian speech, especially Hindi-English code-switching mid-sentence, breaks most Western-trained ASR models. Sarvam costs more per call but gets brand names and mixed-language grocery lists transcribed correctly on the first pass — instead of needing a correction loop.
Streaming feels responsive — the agent starts reasoning before you finish talking — but it's more complex to build and harder to correct mid-sentence if transcription drifts. Batch is simpler and slightly more accurate per utterance, but pause-then-respond reads as laggy for a voice interface.
A smaller model would cut cost per order significantly, but grocery lists have real ambiguity — "the usual milk" needs to resolve against a personalization profile, not a generic SKU. A stronger model gets that right more consistently, which matters more here than shaving cents off an already-cheap operation.
Full autopilot ordering vs a confirm-before-buy step. A direct response to the market signal above — 20% of users don't want to give up control — so the build keeps a human confirm rather than auto-purchasing.
The two tradeoffs that were measurable — streaming vs batch, and model choice — show up in the numbers. Baseline build vs the current streaming build:
Chosen model: Haiku — about 3.7× cheaper at the quality this task needs. Shorter replies cut output tokens further.
Latency sampled from server run logs; cost from the app's own token counters. Not a controlled benchmark.
The software is the brain; the next step is the body. I'm prototyping a small, low-cost kitchen device — a microphone, a speaker, and an ESP32 chip — that sits on your fridge and lets you talk to it while you cook. It streams your voice to the app over Wi-Fi, which does the thinking, and speaks the answer back. It turns your normal fridge into a smart fridge.
Early bench prototype — mic, speaker and an ESP32.