Skip to main content

Introduction

I am Ding Guoliang, an AI Application Engineer. Over the past half year my main line has been an edge voice-agent control platform: Vue + Java control plane for devices, agents, models, knowledge, and team permissions; Python on-device WebSocket speech sessions — listen, recognize, reason, speak, plus tools and on-demand retrieval. Functionally closed and entering launch. I cover the full loop and care about systems you can operate, interrupt cleanly, and degrade safely — not demos only. Desensitized write-up: Voice Agent Platform.

Path

Backend habits still apply: auth, permissions, isolation, failure paths matter as much as models. Those constraints show up in dual-mode config, barge-in, and team data scope.

What I work on

  • Control console and API: devices/agents/models/knowledge, system vs team permissions, config delivery
  • On-device speech service: session path, classic vs end-to-end mutual exclusion, tool dispatch, control-plane sync
  • Frontend: admin UI, permission UX, multi-region business APIs
  • Knowledge: control-plane ops + in-session retrieval

Flagship project

Day-in-the-life, session walkthrough, depth anchors.

Lessons learned

Short retros.

Projects overview

Stack and notes entry.

Introduction

Site entry.

How I think

Production constraints first: felt latency, failure shape, observability. Frameworks can change; control↔session contracts, permission boundaries, and interrupt semantics should not be vague. I write judgments here so they can be revised. Still improving: contract checks, session baselines (first audio / barge-in / tool failure), and where multi-step agents fit in real products.