Listen & respond
Microphone input and speaker output support spoken conversation and interruption.
A desktop companion that listens, sees gestures, and moves. An exploration of what happens when code meets the physical world.
Explore the build ↘What does it take to turn a conversation into a physical action? This robot brings voice, vision, electronics, and motion together in one small system.
The build combines an ESP32-S3 controller, Xiaozhi AI voice firmware, a local vision module, and custom circuit boards. The challenge is making those parts work together—not just getting each part to work on its own.
Microphone input and speaker output support spoken conversation and interruption.
A local AI camera reports gestures. A separate mode controls when they can trigger motion.
Two drive motors, three servos, and a small OLED turn commands into visible responses.
Follow the signal from a spoken command
to a moving motor.
My first custom PCB took three months to build. Getting it to work taught me something I had overlooked: reliable communication starts with reliable power.
I started with an ESP32 development board and got the robot to perform basic movements. But the board was bulky, and loose Dupont jumper connections made the prototype unreliable. I wanted a smaller, more dependable way to connect everything.
After three months of work, I had my first usable PCB—and a major flaw. I had focused on communication circuits without accounting for how the motors would affect the rest of the system. When they ran, microphone operation became unreliable and Wi-Fi grew unstable.
I split the design into two connected boards: one for the 3.3 V controller and low-power peripherals, including the display, and another for the higher-current motor circuitry. I gave the boards separate power feeds, kept a common signal ground, and added a 1000 µF bulk capacitor to help buffer changes in current demand.
The lesson I carried forward: power distribution is part of the system design. I now test the robot with its motors running, not only with its logic circuits powered.
Voice travels over Wi-Fi to the Xiaozhi service. Device tools translate requests into actions. Gesture recognition runs on the SEN0626 module, with the ESP32 checking results before enabling motion.
Implemented control settings: ~150 ms polling, 3 consecutive frames, confidence ≥ 65, gesture 3 allowlist, and a 3-second motor run. Recognition accuracy and gesture mapping still require systematic validation.
Separate controller and interface boards bring the processor, power rails, audio, and motion connections into a compact assembly. Select any image to inspect the design.
Board names follow the supplied schematic and layout assets. PCB renderings are design views; assembly photos appear in the build log.
The enclosure brings together a rotating head, two arms, wheels, and internal electronics. Multiple CAD views show how those parts fit into the robot's compact body.
Physical evidence from the bench.
Design, assemble, debug, repeat.
Custom boards organize the controller and peripheral connections. Assembly makes routing and connector access tangible.
Physical enclosure prototypes connect the CAD design to the realities of mounting, wiring, and servicing the hardware.
Audio, vision, display, and motion come together on the workbench. Integration exposes problems isolated tests cannot.
A prompt sound did not mean the microphone worked. An I²C address did not mean the camera was recognizing gestures. Each signal needed its own test.
A fixed-tone test isolated the speaker path. Microphone sample logs then helped investigate capture, channel selection, and the shared I²S clock configuration.
Camera communication timed out when its external 5 V supply was disconnected. Restoring power made the module detectable at I²C address 0x72.
Gesture mode, repeated-frame confirmation, an allowlist, and timed stopping help prevent a single noisy result from becoming an unintended motor command.
The next step is not adding more features. It is measuring how reliably the existing ones work.
This project uses the open-source Xiaozhi voice firmware, Espressif ESP-IDF, and a DFRobot SEN0626 vision module. The portfolio documents the robot's hardware, integration, and debugging; it does not claim authorship of those underlying platforms.