1. Freeze the dual-mode config contract early
Local defaults help development; control-plane per-device delivery helps operations. If fields diverge, every change becomes two codepaths. Do: separate committed structure from local/secret overrides; share one contract between control plane and device service; update both sides together.2. Speech latency is a product problem
Waiting for a full generation before playback feels broken. Any block in VAD → ASR → LLM → TTS compounds. Do: design for streaming from day one; keep blocking work off the event loop; extend via providers/plugins instead of bloating the connection core.3. Split permission layers
Menu/role auth is not the same as team/resource isolation. Mixing them yields “logged in ⇒ see every tenant”. Do: declare which layer each API uses; enforce isolation in queries/retrieval, not only in the UI.4. Tool failures burn loops
Silent tool failures cause retry storms. Use step caps, duplicate-call detection, wall-clock timeouts, and user-facing degradation.5. Measure before optimizing
Without baselines, model/prompt/retrieval changes are vibes. For speech, track first-audio latency and barge-in success, not only answer quality.6. Observe per connection/device
Without inputs, tool calls, stage latencies, and error codes keyed by session, production becomes restart-only.If I started over
- Freeze the control ↔ device config contract in week one
- Design speech for streaming and degradation from day one
- Decide permission layers before the feature checklist
- Build a small real-session baseline for latency and quality before tuning