AI · Voice · 0→1

A voice AI agent that does your grocery run

You notice you're out of rice and chickpeas mid-cooking — and only order rice, forgetting chickpeas. The job isn't ordering; it's keeping track. So I built a voice-first agent: say what you need, and it finds your usual brands and builds the cart for you.

← Back to portfolio
Live interactive demo · sample catalogue · nothing is really ordered · works best in Chrome. Open demo in new tab ↗
How it works

You talk, the agent does the shopping

1 · Speak

Say "add a litre of milk and some arhar dal" — in English, Hindi, or a mix. It transcribes Indian speech and brand names accurately.

2 · Reason

An LLM agent plans the order: checks your saved address, your usual brands, and your preferences, then searches and picks the right products.

3 · Build the cart

It assembles the basket and reads it back in a natural voice — you review and confirm before anything is bought.

Market validation

The demand signal is real

41%

of professionals said grocery, food & beverage is the number-one category they'd trust an AI agent to complete entirely on their behalf — ahead of travel, electronics, and every other category.

32%

said their biggest worry is that the agent will get it wrong (and 20% don't want to give up control) — exactly why this agent builds the cart but leaves the final confirm to you.

Data based on survey of 1000+ US professionals.

What I built

Prototype and business case, end to end

A 0→1 build I did solo — both the working product and the commercial thinking behind it:

The product

Voice-to-cart agent: speech-to-text → an LLM tool-calling agent (Model Context Protocol) → streaming voice, with a personalization profile and multilingual (Hindi/English) support.

The business case

Per-order and per-user unit-economics model, a customer-validation survey, and a competitive teardown — to test whether it's a real, monetizable product.

The hardware path

Spec'd a low-cost kitchen device (ESP32 + mic + speaker) that turns the phone into the "brain" — the route from software demo to physical product.

Next.js Claude (LLM agent) Model Context Protocol Sarvam (Indian-language STT/TTS) Swiggy Instamart API Streaming voice

Same agent architecture family (Claude + MCP) used by conversational-checkout products in agentic commerce today.

Where it fits

Where this sits in the agentic commerce stack

Voice input, an MCP-based tool-calling agent, and a live commerce API doing the checkout — that's the shape most conversational commerce agents take today: a reasoning layer that plans the order, and a checkout layer that executes it. The reasoning layer is largely solved — Claude- and GPT-class models with MCP-style tool calling handle intent and product matching well. The harder, less-solved piece is the checkout layer itself, specifically agent-native payment authorization: letting an agent transact within a user's pre-approved limits without a human clicking confirm every time. This project keeps that human confirm step deliberately, since spending-limit infrastructure at the payment-rail level is still early across most markets.

Tradeoffs behind the build

Speech recognition — Sarvam over generic ASR

Indian speech, especially Hindi-English code-switching mid-sentence, breaks most Western-trained ASR models. Sarvam costs more per call but gets brand names and mixed-language grocery lists transcribed correctly on the first pass — instead of needing a correction loop.

Streaming voice over batch transcription

Streaming feels responsive — the agent starts reasoning before you finish talking — but it's more complex to build and harder to correct mid-sentence if transcription drifts. Batch is simpler and slightly more accurate per utterance, but pause-then-respond reads as laggy for a voice interface.

Reasoning model — Claude over a cheaper model

A smaller model would cut cost per order significantly, but grocery lists have real ambiguity — "the usual milk" needs to resolve against a personalization profile, not a generic SKU. A stronger model gets that right more consistently, which matters more here than shaving cents off an already-cheap operation.

Trust vs autonomy

Full autopilot ordering vs a confirm-before-buy step. A direct response to the market signal above — 20% of users don't want to give up control — so the build keeps a human confirm rather than auto-purchasing.

The two tradeoffs that were measurable — streaming vs batch, and model choice — show up in the numbers. Baseline build vs the current streaming build:

Latency — time to first audio
Baseline
6.1 s
Final
0.4 s
Latency — full voice turn
Baseline
~18 s
Final
~6 s
Cost per turn — model choice
Sonnet
₹1.02
Haiku
₹0.27

Chosen model: Haiku — about 3.7× cheaper at the quality this task needs. Shorter replies cut output tokens further.

Latency sampled from server run logs; cost from the app's own token counters. Not a controlled benchmark.

Coming up next

Physical AI — turn any fridge into a smart fridge

The software is the brain; the next step is the body. I'm prototyping a small, low-cost kitchen device — a microphone, a speaker, and an ESP32 chip — that sits on your fridge and lets you talk to it while you cook. It streams your voice to the app over Wi-Fi, which does the thinking, and speaks the answer back. It turns your normal fridge into a smart fridge.

Early hardware prototype — ESP32 board, microphone, speaker and breadboard wiring on a desk

Early bench prototype — mic, speaker and an ESP32.