Skip to main content
These lessons come from closing the loop on an edge voice-agent control platform and preparing for launch — real constraints, not abstract checklists. Desensitized. Flagship: Voice Agent Platform.

1. Freeze the dual-mode config contract early

Local defaults help development; control-plane per-device delivery helps operations. If fields diverge, every change becomes two codepaths. Do: separate committed structure from local/secret overrides; share one contract between control plane and device service; update both sides together.

2. Speech latency is a product problem

Waiting for a full generation before playback feels broken. Any block in VAD → ASR → LLM → TTS compounds. Do: design for streaming from day one; keep blocking work off the event loop; extend via providers/plugins instead of bloating the connection core.

3. Split permission layers

Menu/role auth is not the same as team/resource isolation. Mixing them yields “logged in ⇒ see every tenant”. Do: declare which layer each API uses; enforce isolation in queries/retrieval, not only in the UI.

4. Tool failures burn loops

Silent tool failures cause retry storms. Use step caps, duplicate-call detection, wall-clock timeouts, and user-facing degradation.

5. Measure before optimizing

Without baselines, model/prompt/retrieval changes are vibes. For speech, track first-audio latency and barge-in success, not only answer quality.

6. Observe per connection/device

Without inputs, tool calls, stage latencies, and error codes keyed by session, production becomes restart-only.

If I started over

  • Freeze the control ↔ device config contract in week one
  • Design speech for streaming and degradation from day one
  • Decide permission layers before the feature checklist
  • Build a small real-session baseline for latency and quality before tuning