Give the platform away. Get paid on the outcome. Prove it before you raise on it.
A pilot structure built to make "yes" the easy answer for a university that will never sign a software PO — zero platform fee, zero hardware CapEx by default, Vani only gets paid when a seat actually fills. Every number below traces to what's already built: the real outcome-fee engine in the Sheets ledger, the real per-call costs measured off actual calls, and the real Cloud-vs-On-Prem hardware numbers already scoped in a real on-prem deployment quote.
The offer — what a pilot university actually signs
Same outcome-fee mechanism already built into the Sheets ledger (Pricing_Config, auto-routed by course_fee band) — this isn't a new billing system, it's the existing one with the monthly platform line struck to zero and the per-admit line raised to cover it.
| Line item | Full commercial ("Board") terms — already in Config | Pilot terms — proposed |
|---|---|---|
| Setup fee | ₹5,00,000 one-time | ₹5,00,000 one-time — includes the AGX Orin 64GB box (§03/§04); box stays Kcube Labs' property |
| Monthly platform fee | ₹4,00,000/mo | ₹0 — waived in full for the pilot window |
| Outcome fee — offline / degree (flat band, ₹50k–5L course fee) | ₹15,000 / admit | ₹20,000–₹24,000 / admit |
| Outcome fee — premium (>₹5L course fee) | 7% of course fee | 8–9% of course fee |
| Outcome fee — online / bootcamp (micro band, <₹50k) | 6% commission, tiered ₹350–750/admit | 8–10% commission |
| What triggers a fee at all | Only the incremental admit — above the university's own locked baseline_admits. A no-lift pilot costs the university nothing beyond setup. | |
Every fee-bearing event is logged with a real HMAC-signed webhook and an attribution_rule (any_touch_before_enrollment) — the same ledger this session's terminal dashboard reads from. The university sees the same numbers Vani bills against; there's no black box to dispute.
Build your own pilot
Same "make your own PC" idea — the university's own decision-maker sets the dials, the pilot package and the number they'd actually pay comes out the other side. Runs on the same auto-routing logic already built into Pricing_Config (course fee decides the model), not a separate guess.
This is what decides the pricing model automatically — <₹50K = Micro, ₹50K–5L = Flat, >₹5L = Percent. Same band logic already live in the real Sheet.
Only incremental admits above the university's locked baseline are ever billed — this is the number the outcome fee is charged against. Sets the call/WhatsApp volume below automatically.
No hardware, nothing to install — fastest path to a first call.
Telephony and port-fee lines are illustrative reference rates, not a live CPaaS quote — real providers (Exotel, Ozonetel, Sprinklr, Plivo) price channels/ports very differently; get a real quote before this is a signed number (§05).
The outcome-fee line only ever materializes if admits actually happen — a no-lift pilot pays only the setup fee and whatever it actually used.
Absorbed, not billed to the university — the ₹5/min voice pass-through above already covers it. Shown here for Kcube's own margin visibility, not the university's number.
The math behind "not ₹15K, probably more"
Waiving the platform fee isn't free for Vani — it has to be recovered somewhere, or the pilot loses money by design. Here's the actual arithmetic, not a guess.
active_months (4, the Config default per cycle) = ₹16,00,000 not collected as platform fee over one admissions cycle.baseline_admits 500). A believable incremental lift for one program at one university in one cycle: 40–80 admits.Why under-price it on purpose — decided
A pilot that's obviously still cheaper than the university's own cost-per-admit today is the fastest possible yes. The margin comes back at scale, once fee 2–3 succeed and the platform fee returns for a paying, referenceable client — not on pilot #1.
Vani OS Admissions in a Box — what capacity actually costs
One box, sized honestly. Only one row below is something Vani has actually run in production — the rest are engineering estimates from the same component math, clearly marked, not claims.
| Tier | Hardware | Est. cost | Concurrent calls | Status |
|---|---|---|---|---|
| Pilot Box (until now) | Laptop-class GPU (what actually ran the first live pilot's real calls — GTX 1650 Ti, 4GB VRAM) | ₹80,000–₹1,20,000 equiv. | 2 confirmed real in production · real-load-tested to 5 (2026-09-05) — Whisper STT worst-case climbs 2.9s→5.6s, N=1→N=5; Deepgram STT (same box) holds ≤1.5s at N=5 — see callout below | production-run, load-tested |
| R&D / Demo Box (incoming) | NVIDIA Jetson Orin NX 16GB developer workstation (Waveshare/Yahboom/Seeed-class kit) — 1,024-core Ampere GPU, 32 Tensor Cores, 100 TOPS, 102.4 GB/s memory bandwidth, 10–25W. Full bundle: enclosure + fan, Wi-Fi/BT card, 256GB NVMe SSD pre-installed, power adapter, 7" HDMI touchscreen, RK925 mechanical keyboard. Kcube Labs' own in-house rig — not what ships to a pilot university, that's the AGX Orin 64GB row below. | ₹81,000–₹1,18,000 (real sourced quote, full bundle incl. screen/keyboard) | 2–4 (reasoned estimate, scaled down from the AGX Orin's own math below — not yet tested) | ordered — not yet production-verified |
| AGX Orin 64GB Box — chosen | NVIDIA Jetson AGX Orin 64GB Developer Kit — 2,048-core Ampere GPU, 64 Tensor Cores, 275 TOPS, 204.8 GB/s memory bandwidth, 15–60W. NVIDIA-direct, US marketplace. | $3,499 (≈₹2.9–2.95L hardware) → ₹5,00,000 total setup incl. import/integration (§01/§05) | 4–8 (reasoned estimate — see callout below) | component-benchmark estimate |
| Growth Box (alt.) | RTX 4090 tower — already scoped in a real on-prem hardware quote | ₹4,00,000–₹4,50,000 | 10–12 (estimate, "closer floor" sizing) | already in the Lead-to-Ledger doc |
| Scale | Not a box anymore — hybrid or full cloud (LiveKit SFU + cloud STT/LLM/TTS) | Usage-based | No single-box ceiling | Mode A, already built |
Where the 4–8 estimate for the AGX Orin 64GB actually comes from — estimate
No one has published a benchmark of this exact pipeline (Whisper + Ollama + Kokoro, concurrent calls) on this exact kit — this is a synthesis of real component benchmarks, not a measured figure, and it should be treated as a planning range until it's load-tested.
- Whisper — AGX Orin runs ~3.2x slower per-stream than a desktop RTX 4090 (1.6s vs ~0.5s on a reference clip), still comfortably faster-than-real-time on `faster-whisper`. Likely the tightest bottleneck of the three stages.
- LLM (`llama3.2:1b`-class) — on the weaker Orin Nano, throughput scales ~7.5x from single-user to saturated concurrency, bandwidth-bound not compute-bound. The 64GB AGX Orin has ~2x that memory bandwidth, so this should scale meaningfully better than the Nano numbers.
- Kokoro TTS — no Jetson-specific number found; fast on any real GPU, unlikely to be the bottleneck, but unverified on this hardware.
- Real action item, not just a hardware question: Ollama/llama.cpp is explicitly not the recommended serving path for concurrent clients — vLLM is. Worth switching before the load test, not after, or the test will undersell the hardware.
- The R&D box's 2–4 estimate isn't a separate derivation — the Orin NX 16GB sits at roughly half the AGX Orin 64GB's TOPS and memory bandwidth, so its concurrency headroom is scaled down from the 4–8 estimate above by the same ratio, not re-reasoned from scratch.
The R&D box hasn't taken a real call yet — the 2–4 estimate above still stands until it has. The Lenovo itself, though, now has real load-test data — see the callout right below, not just production experience at 2.
Real load test, 2026-09-05: Whisper is the bottleneck, and there's a cheap fix already in the codebase — verified
Ran a real concurrent-call test against the live Lenovo — lk (LiveKit CLI) driving real audio into real rooms via the actual /api/start-call path, not synthetic load. Full runbook + numbers: vani-voice-stack/docs/load-test-runbook.md.
- Whisper STT (the box's default) degrades hard under concurrency — worst-case time-to-first-transcription climbed from 2.9s (N=1) to 5.6s (N=5), roughly +0.5–1s per additional concurrent call. LLM (
llama3.2:1b) stayed flat (~1–1.5s) throughout — confirms Whisper, not the LLM, is the real ceiling. - Deepgram STT — already a real, working code path in bot.py, just needed its credential wired through (found and fixed live) — swapped in for the same test, same box, same concurrency: worst-case stayed under 1.5s all the way to N=5, essentially flat. Moving STT off the box's CPU fixes the degradation directly, cheaper than assuming only new hardware can.
- Two real bugs found and fixed along the way: Sarvam TTS's
bulbul:v2model had been deprecated server-side (every Sarvam-path call was failing its TTS stage) — bumped tobulbul:v3. Separately,DEEPGRAM_API_KEYwas sitting unused in a `.env` file but never actually wired into the container's environment — same class of gap as the earlierMAX_CONCURRENT_CALLSmiss. Both fixed and deployed same day. - What this doesn't answer yet: none of this ran on the AGX Orin box — this was the Lenovo only. The bandwidth/TOPS-ratio reasoning above for the Orin's 4–8 estimate still stands as reasoned, not measured, for that hardware. But it does mean the "buy bigger hardware" lever isn't the only one on the table — a cloud STT swap is a real, already-built, cheaper alternative worth weighing first.
Does the on-prem box need its own GCP hosting cost? Reasoned no — not yet tested — estimate
Today's real call transport runs through a LiveKit SFU self-hosted on the Lenovo itself (via a Cloudflare Tunnel + TURN-over-TLS — GCP has been fully decommissioned as of 2026-09-07). The AGX Orin box runs real Linux (JetPack/Ubuntu) too — the same self-hosting pattern already proven on the Lenovo should carry over to the box without a separate cloud hop.
So: no separate recurring cloud hosting line for the on-prem box — it should be self-contained except for a domain + TLS cert (near-zero cost, same pattern already proven on the Lenovo). What's not free is time: the university's own network needs to open/forward the same ports already proven in that self-hosted setup (TCP 7880-7881 signaling, UDP 50000-50100 media, or the Cloudflare Tunnel equivalent) before external calls can reach the box — a real site-provisioning step, already folded into §07's "weeks not days" On-Prem lead time, not an extra cost.
This has never actually been run — real LiveKit, self-hosted, on real Jetson Linux, taking a real call. Treat it as a reasoned architectural inference until it's tested once, the same way the 4–8 concurrency estimate above is.
Who pays for the box
The whole pitch is "no capital expense." So the default has to actually deliver that — the box being available for purchase isn't the same as the university being asked to buy it.
Vani/Kcube Labs owns the box
- UpfrontKcube Labs procures the AGX Orin 64GB kit direct from NVIDIA (US, $3,499) and ships/installs it. University pays ₹0 hardware CapEx.
- OwnershipThe box is and remains Kcube Labs' own property, for the life of the pilot and after — not transferred, not depreciated on the university's books. The ₹5,00,000 setup fee (§01) is a setup-and-hardware-provisioning service charge, not a sale.
- Recovered viaThe ₹5,00,000 one-time setup fee, not the per-admit outcome fee — keeps §02's outcome-fee math clean and lets the setup fee stand on its own as "what it costs to stand this up."
- Who maintains itKcube Labs, remotely — same as the software layer already is.
- Best fitEvery pilot by default. This is the actual "skin in the game" move: Kcube Labs is visibly carrying the hardware risk, not the university — and can reclaim/redeploy the box if a pilot doesn't convert.
University owns the box
- What this would look likeSame ₹5,00,000 setup cost either way — installed at the university's own site either way. The only thing that changes is who holds title to the box afterward: here, the university, not Kcube Labs.
- Why it's not the defaultDefeats the actual point of this plan — the CapEx-avoidance pitch is the reason a decision-maker can say yes in one meeting. Keeping the box as Kcube Labs' property is what makes "skin in the game" a real, visible claim in the room, not just a line in the deck.
- When it might still come upA university that's already converted from pilot to paying client and specifically wants full infrastructure ownership going forward — a second-conversation option, never the opener.
| Setup fee breakdown | Amount | Note |
|---|---|---|
| AGX Orin 64GB Developer Kit | ≈₹2,90,000–₹2,95,000 | $3,499 at current INR/USD — verify the exact rate at time of purchase, this is not a live quote. |
| Import, customs, logistics | Balance of ₹5,00,000 | Real cost, not yet itemized — India import duty on this HS classification needs an actual customs check before this number is final, not assumed. |
| Integration & setup labor | included above | CDL/persona setup for the institution's own programs, CRM/LMS connectors, phone number + WhatsApp provisioning, Deepgram STT account/credential setup (§04's real load test) — Kcube Labs does this directly, not left to the university, so the box performs at its tested-good numbers from day one. |
| Total setup fee, university-facing | ₹5,00,000 | One-time. Box remains Kcube Labs' property (above). |
| Ongoing service cost — either path | Rate | Who's billed, and how |
|---|---|---|
| AI voice minutes | ₹5/min | voice_ai_rate_per_min_inr — already in Config as the "Voice Cost Transparency Clause," pass-through either way. |
| WhatsApp messages | ₹2/msg (Config default) — real BSP cost ₹0.11–0.86/msg | University's own Meta Business/BSP account in both Cloud and On-Prem modes — never bundled behind the platform fee, confirmed in the Lead-to-Ledger doc. |
| Telephony (On-Prem box only) | ₹0.60–1.20/min | University's own Plivo/Exotel account, direct — same number already scoped in that on-prem hardware quote. |
| CPaaS concurrent-channel/port fee | Illustrative ~₹2,000/channel/mo — not a live quote | A real, separate cost some providers bill on top of per-minute usage (Ozonetel-class: ~$25/channel/mo). Others (Exotel's published plans) bundle unlimited channels instead. Depends entirely on which CPaaS the university's telephony account uses — needs a real vendor conversation before this is a number in an agreement (§02's configurator sizes it off estimated peak concurrency, for reference only). |
| Deepgram STT (§04's real load-test fix) | $0.0077/min standard (≈₹0.64/min) — a $0.0048/min promotional rate exists today but isn't guaranteed to last, so this plan prices at the standard rate | Vani/Kcube Labs, absorbed — not billed to the university. The ₹5/min AI-voice pass-through above is unchanged either way; this is what it costs Vani to actually hit the tested-good latency numbers in §04, set up directly by Kcube Labs as part of onboarding, not left for the university to configure. |
Sequencing — pilot, prove, then raise
The first live test of exactly this shape of deal is already underway — this formalizes the pattern for the 2nd and 3rd, not a new motion from zero.
Pick 2–3 pilot profiles, not 10
One offline-degree institution (already underway) plus one online/bootcamp partner to actually exercise the percent-commission model, and one more offline institution as the second reference point. Three is enough to prove the pattern repeats; ten is a sales motion this plan isn't funded for yet.
Run it for real, on the real ledger
Zero platform fee, live on the pilot terms above. Every lead, handoff, and admit already flows through the same webhook-signed Event_Log and terminal dashboard built and verified this week — the proof engine isn't something to build later, it's already watching.
This is also where the Box's real concurrency ceiling gets found under real, not synthetic, load — feed that number back into §03 before quoting a 4th university.
Raise on proof, not projection
Real ROI multiple, real cost-per-admit, real incremental-admit count, from the same dashboard an investor can be shown live — not a slide someone built for the deck. This is the actual argument this whole plan exists to earn the right to make.
Convert the pilots, don't abandon them
Pilot universities move to full commercial terms — at an early-adopter rate, not the sticker price — in exchange for staying on as public references. Burning the first believers to hit a clean price sheet costs more in the next 5 deals than it saves in this one.
What's still open
Before this goes in front of a second university — open risk
- Load test past 2 concurrent calls — done on the Lenovo (2026-09-05, real test with real audio, up to N=5 — see §04's callout), not done on the AGX Orin/Growth Box tiers themselves. Those numbers are still estimates until real calls on that actual hardware prove them.
- Attribution disputes —
any_touch_before_enrollmentis the rule, but a university disputing "was this admit really ours" needs a clear, pre-agreed resolution path before it's a real disagreement, not during one. - Regulatory read — outcome-fee-based lead-to-admission models can draw UGC/regulatory attention in India depending on structure. Worth a real legal read before this is a signed pilot agreement, not just a plan.
- Hardware lead time — the Lead-to-Ledger doc already flags On-Prem as "weeks, not days" to go live. Kcube Labs financing/owning the box (§05's default) makes that lead time Kcube's problem to manage across 3 simultaneous pilots, not one. Now also includes getting the university's own IT team to open/forward the box's LiveKit ports (§04) — a real coordination step, not just shipping.
- Customs/import duty on the AGX Orin kit — §05's ₹5L breakdown doesn't itemize this yet; needs a real check before it's a number in a signed agreement.
- Self-hosted LiveKit on the box, untested — §04's "no separate cloud cost" claim is a reasoned inference from the real self-hosted setup already proven on the Lenovo (Cloudflare Tunnel + TURN-over-TLS), not a proven deployment on this specific hardware. Load-test it on real Jetson hardware before it's a claim made to a university, not just a plan.
- CPaaS channel/port pricing, unconfirmed — §02/§05's ~₹2,000/channel/month is illustrative (one real reference point, not this pilot's actual provider). Get a real quote from whichever CPaaS the pilot university's Plivo/Exotel-class account uses before this is a number in an agreement.
Reality check — has anyone actually done this
Real market research, done before shipping this plan, not a hunch. Every piece of this model already exists somewhere — the specific combination doesn't appear to, yet.
What's already proven, at scale — decided
- Outcome-based AI-agent pricing is now a real industry pattern. Zendesk charges $1.50–2.00 per automated resolution, Intercom's Fin $0.99/outcome, HubSpot $1/qualified lead. Nothing found prices this specifically for university admissions yet — but the pricing shape itself is no longer unusual or unproven for AI agents in general.
- Full skin-in-the-game revenue share is a decades-old, multi-billion-dollar category in higher ed. Online Program Management (2U, Academic Partnerships, Risepoint, and others) routinely takes 20–94% of tuition revenue — commonly 50–60% — in exchange for upfront marketing/operations investment. Structurally, that's the exact same trade this plan proposes: a vendor carries the cost, gets paid from the outcome — just applied to a whole degree program instead of one admissions funnel, at a vastly higher cut.
- Indian universities already buy performance-priced admissions services today. Study-abroad agents earn 10–20% commission per successfully placed student, and domestic lead-gen agencies already sell admissions leads at a real cost-per-lead (₹168–500 seen in a published case study). This pricing shape isn't a hard sell to an Indian decision-maker — they already work this way with human vendors.
| Direct AI-voice-admissions competitor (India) | What they actually sell | Pricing model |
|---|---|---|
| Meritto / Mio AI Voice (NoPaperForms) | AI voice agent inside a full admissions CRM — live with real universities across 6 states in the 2026 cycle | Software subscription + consumption — not outcome-priced |
| Caller Digital | Voice AI for admission funnels, India-focused | Subscription (platform-fee shaped) |
| Edesy | AI voice agent for college admission calls | Subscription (platform-fee shaped) |
| VaniAgent (Solvea Technologies) | General-purpose India voice AI; lists admissions counseling as one use case | ₹3–8/min + platform fee |
None of the four found here price on a per-admit or outcome basis — every one is a subscription or consumption meter, the same shape as this plan's own full commercial terms (§01). The AI-automation side of this market is real and getting crowded; the outcome-fee side of it doesn't appear to be claimed yet.
The real cautionary tale — and why it shouldn't repeat here — open risk
OPM revenue-share deals are the closest large-scale precedent for "pay on outcome" in higher ed — and they've drawn real, documented criticism. Four failure modes, and what's structurally different in this plan:
- Revenue extraction — OPMs profit more as tuition rises, so contracts have incentivized fee hikes. This plan's fee is fixed per admit or percent of a course fee the university sets independently — Vani has no lever to push tuition up, and no reason to.
- Hidden vendor involvement — students often can't tell OPM staff from university staff. Vani OS Admissions is a calling/WhatsApp assistant, not a person impersonating one — disclosure is a design constraint already, not an afterthought.
- Lock-in contracts — 6+ month exit windows and auto-renewal traps are common in OPM deals. This plan is explicitly a short, terminable pilot (§06) — no multi-year lock-in proposed at this stage.
- A regulatory loophole being actively challenged — US law (HEA §487(a)(20)) bans paying enrollment-based commission to recruiters outright, with a "bundled services" exception OPMs rely on that reform advocates want closed, and real enforcement exists (a $1.3M settlement in 2026). This doesn't apply to Vani OS Admissions' India-only pilots today — but it's a hard rule to know about before ever pricing a US deployment this way.
Verdict — decided
Nothing found combines AI-automated admissions outreach with true outcome-based pricing as a packaged offer. Every adjacent precedent does one half: OPMs and study-abroad agents do outcome-based pricing without the AI automation; Meritto, Caller Digital, Edesy, and VaniAgent do the AI automation without outcome-based pricing. That gap looks real, not an oversight — worth treating as the actual pitch, not just a nice-to-have.