Back to Selected Work
Case Study2025
AI/MLSpeech IntelligenceMulti-Tenant SaaSData Engineering

Hana Voice AI

Audio-first intelligence that remembers who said what, and when.

Role & Services

System Architecture, ML Pipeline Engineering, Full-Stack Development

Timeline & Year

2025

Deployment Status

Production Architecture / Enterprise Pilot

Hana Voice AI Hero Overview
Hana Voice AI — System Overview1920 × 1080 (16:9)
[ At-A-Glance Specification ]

Target Industry

Enterprise meetings, compliance, knowledge management; India-first workplaces

Core Audience

Enterprise executives, compliance officers, engineering leadership

Stack Foundations

Next.js 16 (App Router), PostgreSQL, Clerk (Multi-Tenant Organizations)

Key Metric

9 Stages

Asynchronous Pipeline

[ Operational Scoping ]

The Problem & The Applied Solution

The Operational Challenge

Fragmented tools, coordination fatigue, and unverified data.

Teams in modern Indian and multilingual workplaces switch fluidly between English, Hindi, and Hinglish mid-sentence. Conventional monolingual tools produce gibberish transcriptions, reset speaker identifiers on every new call, lose critical operational commitments, and silently propagate contradictory statements across corporate wikis.

Impact: Coordination overhead, high human-error rate, fragmented context.
The Engineering Solution

Autonomous agents, verified state machines, and grounded execution.

A resilient 9-stage asynchronous pipeline: UPLOADED → VALIDATING → PREPROCESSING → DIARIZING_ASR → IDENTITY → UNDERSTANDING → MEMORY → INDEXING → READY. Raw audio undergoes normalization (16 kHz mono PCM WAV, -20 dBFS RMS) and VAD silence stripping, followed by dynamic language routing (en / hi / hi-en). Millisecond word-aligned ASR pairs with Pyannote diarization and 512-dimensional voiceprints. Probabilistic Bayesian identity resolution fuses acoustic and organizational signals to link speech to canonical person entities, updating bi-temporal fact graphs, commitment state machines, and grounded RAG search.

Outcome: Unified operational dashboard, deterministic type safety, automated verification.
[ System Design & Topology ]

9-Stage Asynchronous Audio Intelligence Pipeline

The audio intelligence pipeline decouples compute-intensive signal processing and machine learning inference from user-facing APIs. Non-blocking state machines process uploaded audio through nine verifiable validation, diarization, identity, and memory ingestion phases.

Signal & ML Architecture

9-Stage Asynchronous Pipeline & Identity Engine

Non-Blocking State MachineSHA-256 Deduplicated

[ 9 Asynchronous State Stages : Click to inspect transition criteria & invariants ]

STAGE 04

DIARIZING_ASR

Multilingual code-switching ASR & Pyannote diarization
  • Checkpoint 01Dynamic language routing: en / hi / hi-en mid-sentence routing
  • Checkpoint 02Whisper prompt conditioning with enterprise vocabulary biasing
  • Checkpoint 03Pyannote spectral clustering generating 512-dimensional voiceprints
  • Checkpoint 04Millisecond-level word timestamp alignment and confidence scoring

Pipeline Dataflow: Signal → Embedding → Identity → Memory → RAG

01. Audio Ingest
Raw Multi-Format Audio→SHA-256 Tenant Dedup→16 kHz Mono PCM WAV→-20 dBFS RMS→VAD Silence Stripping
02. Speech & Acoustic
Dynamic Language Router (en / hi / hi-en)→Whisper Alignment (Word Timestamps)||Pyannote Diarization (1.5s windows / 0.75s overlap)→512-d Voiceprints
03. Identity Fusion
Voiceprint (0.45)+Calendar Metadata (0.25)+Semantic Context (0.15)+Graph Co-occurrence (0.15)→Canonical Person Mapping
04. Memory Engines

Threads Engine

Jaccard overlap & topic decay

Bi-Temporal Memory

Assertion & supersession

Commitment Tracker

Promises with audio proof

Contradiction Engine

Dual-audio conflict cards

Knowledge Graph

kg_nodes & kg_edges JSONB

05. Grounded RAG
Hierarchical Scope Routing→Full-Text + pgvector→Grounded Answers with Playable Audio Citations
Multi-Tenant Security: Clerk org_id row-level database isolationLocal / On-Premise Capable with PGlite Embedded Fallback
01

Signal Normalization: SHA-256 tenant deduplication, 16 kHz mono conversion, and -20 dBFS leveling before any model execution.

02

Code-Switching Routing: Dynamic routing dynamically dispatches between English, Hindi, and Hinglish vocabulary models with prompt conditioning.

03

Bayesian Identity Fusion: Weighted scoring (voice 0.45, metadata 0.25, context 0.15, co-occurrence 0.15) maps utterances to persistent canonical profiles.

04

Bi-Temporal Fact Ledger: Tracks assertion time, system ingestion time, and validity intervals with immutable supersession chains.

05

Grounded Audio Citations: Every extracted decision, action item, or query response links directly to playable millisecond-accurate audio proof.

Architectural Differentiation

Hana Voice AI vs. Generic Transcription Bots

Why commodity meeting recorders fail in multilingual, high-accountability enterprise settings.

Operational DimensionHana Voice AI ArchitectureGeneric Transcription Bots
Speaker Identity
✓Persistent 512-d voiceprints recognizing individuals across meetings.
✕Generic speaker labels (Speaker 1, 2) that reset on every call.
Language Support
✓Dynamic mid-sentence English / Hindi / Hinglish code-switching.
✕Monolingual models that fail on multilingual phrases.
Organizational Memory
✓Bi-temporal facts with validity intervals and non-destructive supersession.
✕Ephemeral transcripts and static, disjointed text notes.
Conflict Detection
✓Automated contradiction engine surfacing dual audio clips for verification.
✕Silent accumulation of opposing statements across documents.
Action Items & Promises
✓State machine tracking lifecycle with millisecond-exact audio proof.
✕Unverified markdown bullet lists without ownership accountability.
Information Retrieval
✓Grounded answers with playable millisecond audio citations.
✕Unverified summaries with potential hallucination risk.
Deployment & Privacy
✓Local / on-premise capable with strict row-level tenant isolation.
✕Public cloud lock-in with shared database pools.
Standard: Grounded answers with audio citationsZero cloud lock-in architecture
[ Core Functional Capabilities ]

Engineered Capabilities & Workflows

Each capability module represents a discrete, type-safe subsystem designed for resilience and performance.

  • ✦SHA-256 tenant-isolated hash deduplication preventing redundant processing of identical audio files.
  • ✦Lossless FFmpeg audio normalization to standardized 16 kHz mono PCM WAV format.
  • ✦RMS dynamic leveling calibrated to -20 dBFS for optimal acoustic model sensitivity.
  • ✦Voice Activity Detection (VAD) stripping silence and non-speech artifacts before pipeline ingestion.
  • ✦Resilient non-blocking job state machine with automated retry policies and dead-letter queues.
[ Production Infrastructure ]

Full-Stack Architecture & Technology Stack

Strictly vetted dependencies ensuring type safety, low latency, and zero vendor lock-in.

Layer 01

Core & Framework

Next.js 16 (App Router)React 19TypeScript Strict Mode
Layer 02

Data & Storage

PostgreSQLDrizzle ORMPGlite (Embedded Local Database)pgvector
Layer 03

Multi-Tenancy & Auth

Clerk (Multi-Tenant Organizations)Row-Level Security (RLS)Role-Based Access Control
Layer 04

Audio & ML Pipeline

FFmpegWebAudio APIWhisper / Faster-WhisperPyannote.audio
Layer 05

UI & Internationalization

Tailwind CSS v4Radix UILucide Iconsnext-intl (EN / HI / FR)
Layer 06

Quality & Observability

VitestPlaywrightStorybook 10ESLint & KnipSentry & LogTape
[ Quantifiable Impact ]

Key Engineering Highlights & Outcomes

9 Stages

Asynchronous Pipeline

Zero-blocking processing pipeline from raw WAV upload to indexed knowledge graph.

512-dim

Voiceprint Embeddings

Persistent acoustic voiceprint clustering that resolves speaker identity across meetings.

4 Signals

Bayesian Identity Fusion

Fuses acoustic similarity, meeting metadata, semantic context, and co-occurrence graphs.

100 ms

Exact Audio Citations

Every commitment, contradiction, and factual claim links directly to timestamped audio proof.

[ Product Interface Showcase ]

Visual Workspace & Operational Interfaces

Inspect primary screens, interactive cards, and live operational consoles. Click any image to expand full-resolution view.

Hana Voice AI main workspace showing processed meetings and participant voiceprints
Click to Expand Lightbox

Multi-Tenant Audio Intelligence Command Center

1920 × 1080 (16:9)

Primary meeting view with waveform player, diarization timeline, participant identity pills, and live transcript follower.

Operational telemetry view showing real-time progression through the 9 pipeline stages
Click to Expand Lightbox

9-Stage Asynchronous Pipeline Telemetry

1600 × 1000 (16:10)

Pipeline status monitor showing stage timestamps, deduplication confirmation, language routing score, and diarization status.

Audio player showing color-coded speaker segments and synchronized Hinglish transcript
Click to Expand Lightbox

Multilingual Code-Switching Waveform Player

1600 × 1000 (16:10)

Waveform player demonstrating 0.75x-2.0x playback, speaker turn markers, and aligned Hindi/English words.

Side-by-side contradiction card comparing statements from two meetings with audio proof
Click to Expand Lightbox

Contradiction Engine Dual-Audio Verification

1600 × 1000 (16:10)

Dual audio card highlighting opposing budget or timeline claims made across different calls, with resolution actions.

Interactive node-link graph showing connections between speakers, decisions, and projects
Click to Expand Lightbox

Interactive Postgres Knowledge Graph Traversal

1600 × 1000 (16:10)

Relational graph view rendering person nodes, meeting nodes, and commitment edges with multi-hop drill-down.

[ Engineering Partnership ]

Have a similar engineering challenge?

Whether architecting multi-agent generative systems, private on-premise AI, or speech intelligence pipelines — we engineer solutions that perform in production.