Skip to main content
Desensitized flagship case study: an edge voice-agent control platform I helped build over the past half year and more. The goal was a shippable loop — operate in the cloud, run real-time voice agents on the edge — not three disconnected demos.
No company names, product brands, customers, or internal details. Architecture roles and engineering trade-offs only.

Problem and goals

Edge voice needs low latency and stable sessions; the cloud must manage devices, agents, models, knowledge bases, and team permissions. Without a shared contract you get:
  • Config that cannot be delivered per device
  • Assistants stuck in single-turn Q&A without tools or retrieval
  • Fuzzy tenancy boundaries that block SaaS-style operations
We needed a control plane + session plane platform: operate in the console, run speech on the device service, with a clear deployable contract.

System overview

Control UI

Console for devices, agents, models, knowledge, and permissions.

Control API

Business/config backend: auth, team isolation, delivery contracts.

Session plane

Real-time speech and tool orchestration driven by config.

What I helped build

Framed as contributed to (not sole ownership of the full stack):
  • Console and API: device/agent/model/knowledge domains, team permissions and auth, config-delivery contracts to the device service
  • On-device speech service: session path and message routing, ASR/TTS (and end-to-end speech) integration, tool/plugin/MCP dispatch, dual-mode config with the control plane
  • Admin web: management UI, multi-environment API wiring, permission-aware interactions
  • Retrieval: knowledge/vector search (including image-vector scenarios) into the platform
Deeper themes: Agents, RAG, Multimodal, Infrastructure.

Key design choices

Dual-mode configuration

  • Local mode: defaults + local overrides for development
  • Control-plane mode: pull per-device config at connect time; secrets stay out of committed defaults

Real-time session pipeline

Classic path: audio → VAD → ASR → intent/tools → LLM → TTS → playback. An end-to-end speech path is optional. Keep the connection core thin; extend via providers, plugins, and message handlers.

Tools and extensions

Unify server plugins, MCP, and device capabilities behind one dispatcher so “can talk” becomes “can act”. Schemas, errors, and timeouts are part of production readiness.

Knowledge and vectors

Control plane manages knowledge and models; retrieval uses a vector store (e.g. Qdrant) for documents/images before generation. Multi-tenant filtering must be explicit at retrieval time.

Permission layers

  • System permissions (menus/roles)
  • Team/resource permissions (business data isolation)
Every API should declare which layer it belongs to.

Launch-stage trade-offs

With the loop closed and launch underway, priorities tend to be:
  1. Session stability and latency — speech is sensitive to first token and barge-in; keep streaming and async/thread-pool boundaries clear
  2. Graceful degradation — ASR/LLM/TTS failures must not kill the connection
  3. Config and secret hygiene — no secrets in committed defaults; contract changes stay bilateral
  4. Observability — logs keyed by connection/device, or production becomes restart-only
More in Lessons Learned.

Stack (aligned with this case)


Projects overview

Projects section entry and stack notes.

Agents

Agent and tool-use deep-dives.

Multimodal

Speech and vision pipelines.

Lessons learned

Launch-stage pitfalls.