No company names, product brands, customers, or internal details. Architecture roles and engineering trade-offs only.
Problem and goals
Edge voice needs low latency and stable sessions; the cloud must manage devices, agents, models, knowledge bases, and team permissions. Without a shared contract you get:- Config that cannot be delivered per device
- Assistants stuck in single-turn Q&A without tools or retrieval
- Fuzzy tenancy boundaries that block SaaS-style operations
System overview
Control UI
Console for devices, agents, models, knowledge, and permissions.
Control API
Business/config backend: auth, team isolation, delivery contracts.
Session plane
Real-time speech and tool orchestration driven by config.
What I helped build
Framed as contributed to (not sole ownership of the full stack):- Console and API: device/agent/model/knowledge domains, team permissions and auth, config-delivery contracts to the device service
- On-device speech service: session path and message routing, ASR/TTS (and end-to-end speech) integration, tool/plugin/MCP dispatch, dual-mode config with the control plane
- Admin web: management UI, multi-environment API wiring, permission-aware interactions
- Retrieval: knowledge/vector search (including image-vector scenarios) into the platform
Key design choices
Dual-mode configuration
- Local mode: defaults + local overrides for development
- Control-plane mode: pull per-device config at connect time; secrets stay out of committed defaults
Real-time session pipeline
Classic path:audio → VAD → ASR → intent/tools → LLM → TTS → playback. An end-to-end speech path is optional. Keep the connection core thin; extend via providers, plugins, and message handlers.
Tools and extensions
Unify server plugins, MCP, and device capabilities behind one dispatcher so “can talk” becomes “can act”. Schemas, errors, and timeouts are part of production readiness.Knowledge and vectors
Control plane manages knowledge and models; retrieval uses a vector store (e.g. Qdrant) for documents/images before generation. Multi-tenant filtering must be explicit at retrieval time.Permission layers
- System permissions (menus/roles)
- Team/resource permissions (business data isolation)
Launch-stage trade-offs
With the loop closed and launch underway, priorities tend to be:- Session stability and latency — speech is sensitive to first token and barge-in; keep streaming and async/thread-pool boundaries clear
- Graceful degradation — ASR/LLM/TTS failures must not kill the connection
- Config and secret hygiene — no secrets in committed defaults; contract changes stay bilateral
- Observability — logs keyed by connection/device, or production becomes restart-only
Stack (aligned with this case)
Read next
Projects overview
Projects section entry and stack notes.
Agents
Agent and tool-use deep-dives.
Multimodal
Speech and vision pipelines.
Lessons learned
Launch-stage pitfalls.