Build a Talking AI Desk Robot You Flash From a Web Page

by robolynk in Circuits > Microcontrollers

94 Views, 3 Favorites, 0 Comments

Build a Talking AI Desk Robot You Flash From a Web Page

01-cover-assembled.jpg

Most "build an AI assistant" projects stall in the same place: install a

toolchain, pick a board definition, chase library versions, and get through a

wall of red text before you hear a single word.


This is the opposite. An off-the-shelf 2.8" ESP32-S3 board, a 630 mAh battery, a

printed case, and firmware that flashes straight from a browser tab. Press the

screen, ask a question, let go. It answers out loud with an animated face, sets

alarms, and turns your lights off when you say good night.


**About ten minutes of hands-on work.** Flashing takes 70 seconds and assembly a

few minutes. The only long part is the printer — roughly 1 hour 20 minutes for

all four case parts, and you can flash and set the robot up while it runs.


There is no custom PCB and no wiring at all. The only time a soldering iron comes

out is to press four heat-set inserts.


Watch the whole build, start to finish (5 min): https://youtu.be/vgklh-3mNh4


Everything — firmware, printable files, parts list and the one-click flasher —

is free: https://varozhan3212.github.io/robolynk-light/


Supplies

02-parts-laid-out.jpg
07-parts-wide.jpg

**Electronics**

- ES3C28P (or ES3N28P) ESP32-S3 board × 1 — 2.8" 240×320 ILI9341 display,

capacitive touch, ES8311 audio codec, speaker amp, microSD slot, USB-C

- LiPo battery, 3.7 V 630 mAh, JST-PH, **with protection circuit** × 1

- 8 Ω speaker, small × 1

- microSD card × 1 (optional — only for the larger face animation)

- GPS module, UART 3.3 V × 1 (optional — the only part that needs wiring)


**Hardware**

- M2 heat-set inserts × 4

- M2 screws, 6–8 mm × 4


**Tools**

- 3D printer (PETG or PLA)

- Soldering iron — only for pressing the inserts

- Hot glue gun

- USB-C cable and a computer running Chrome or Edge

⚠️ **Match the board model number.** Other ESP32-S3 boards look identical but use

a different display controller or have no audio codec, and the firmware will not

run on them. This is the single most common way to end up with a dead build.

Start the Print First

02-parts-laid-out.jpg
07-parts-wide.jpg

PHOTO: B/02-printed-shells.jpg

PHOTO: A/05-red-and-black-shells.jpg


Printing is the only step you wait on, so start it before anything else. You can

flash the board and set it up while the printer runs.


Four files, all on one plate, **no supports**:


| File | Qty | Notes |

|---|---|---|

| front_shell_v22.stl | 1 | Holds the display; the bezel sits flush |

| back_cover_v22.stl | 1 | USB-C and microSD stay accessible |

| clips_v22.stl | 1 plate | Internal clips |

| buttons_v22.stl | 1 plate | Caps for the reset and boot buttons |


**Settings:** PETG or PLA · 0.2 mm layers · 20% infill · **no supports** ·

orientation as the files load — no reorienting needed.


About **1 h 20 min** for all four parts. Printed on a Bambu Lab with stock

profiles. The clips and buttons are small and fit alongside either shell.


Files: https://github.com/varozhan3212/robolynk-light/tree/main/case


Flash the Firmware From a Browser (70 Seconds)

09-flashpage.png

No toolchain, no IDE, no libraries.


1. Plug the board into your computer with USB-C.

2. Open the flasher in **Chrome or Edge** (Safari and Firefox cannot do this —

they do not support WebSerial): https://varozhan3212.github.io/robolynk-light/

3. Press **Connect**, pick the serial port, press **Install**.

4. Wait about 70 seconds.


That is the whole process. The firmware is a single image containing the runtime,

the application and all its files — there is no second step and nothing to copy

afterwards.


When it finishes, the board restarts and shows the splash screen.


Press in the Heat-Set Inserts

Set the soldering iron **low** and press the four M2 inserts into the front

shell, holding the iron straight down so the insert goes in square.


Go slowly and let each one cool before moving to the next. Rushing tilts them,

and a tilted insert pulls the screw off-axis so the back cover never sits flat.


Fit the Board

03-board-in-case.jpg

Drop the board in **display-first** and seat the bezel so it sits flush.

Then **hot-glue the board at the corners.** Nothing else stops it shifting, and a

loose board strains the USB-C port every single time you plug in — that is how

these get killed. Keep glue off the connectors and clear of where the battery

will sit.

Battery (Check Polarity First)

04-battery-and-board.jpg


🚨 **Check the battery polarity before you plug it in.** JST-PH connectors are

not standardised between suppliers, and a reversed cell will damage the board. If

your battery's housing is wired backwards, move the pins in the housing rather

than forcing it in.


Once you are sure, connect it and tuck the cell in so the shell does not pinch it

when it closes.

Close It Up

04-battery-and-board.jpg
05-red-and-black-shells.jpg

1. Push the button caps in from the outside.

2. Fit the back cover.

3. Four M2 screws.


USB-C and the microSD slot stay accessible with the case closed, so you never

have to open it again — including for firmware updates, which happen over Wi-Fi.

Wi-Fi Setup From Your Phone

There is no keyboard and no hard-coded credentials in the firmware.


1. Power on. The robot shows **TAP TO BEGIN**.

2. It raises its own hotspot called **RoboLynk-XXXX**.

3. Join that hotspot from your phone and open **192.168.4.1**.

4. Pick your network from the list and enter the password.


⚠️ **2.4 GHz only.** The board has no 5 GHz radio, so a 5 GHz-only network will

not appear in the list.


Pair It and Say Something


1. Create an account at https://robolynk.cloud

2. On the robot, go to **Settings → Info** and read the **six-character setup

code**.

3. Add the robot with that code.


Now hold the screen, ask a question, and let go.


It remembers the last six messages, so follow-up questions work — ask "who wrote

it?" straight after "what's Dune about?" and it knows what you mean.

What Else It Does


- **An animated face** while it speaks.

- **Alarms and timers** by voice.

- **Philips Hue control** — "turn on the lights", "dim the bedroom",

"lights to 50%".

- **A built-in light and an SOS signal that work with no network at all.** Tap

the top-left of the screen for the white LED; hold it three seconds for an SOS

pattern. Neither needs Wi-Fi or a server.

- **It updates itself** over Wi-Fi.


Reminders and alarms here are a convenience feature. Do not rely on them for

anything that matters medically.

Three Bugs That Looked Like Hardware Limits

Worth reading if you are building anything that captures audio on a

microcontroller. All three presented as "the chip isn't fast enough". None were.


**1. The first word of every question was missing.** People asked "why is the 4th

of July important" and the transcript came back "July is important". The cause

was ordering: the code updated the screen and ran a garbage collection *before*

switching the microphone on. People start speaking the instant the screen says

LISTEN, so the opening word was never recorded. The microphone now starts first

and the housekeeping happens behind it.


**2. The capture loop was sipping from the buffer.** It read one 30 ms slice of

audio per network send, and each send costs ~110 ms — so it captured barely a

quarter of real time and the rest piled up until the buffer overflowed. Reading

until the buffer is actually empty, and sending larger batches, took capture from

27% to 85% of real time. The detail that cost the most: **a blocking read never

reports that it has run dry — it just waits.** The loop looked healthy the whole

time it was losing audio.


**3. Streamed speech had nowhere to wait.** Playing straight from the network

into a 170 ms hardware buffer means any late packet is an audible gap. Holding

~600 ms back before starting playback, into a larger buffer, removed it entirely.


When embedded audio misbehaves, suspect your own ordering and buffering long

before you blame the silicon.

Things Worth Knowing

- **It needs a server.** The speech and language work does not happen on the

board. This project is the device half.

- **The face animation is the memory budget.** Frames live in RAM, and the loader

reserves headroom for the TLS handshake and audio buffers before deciding how

many frames fit. Take that headroom away and it boots to a white screen.

- **A microSD card on a bit-banged SPI bus is slow** — about 93 KB/s here, so a

multi-megabyte animation takes over a minute to load at boot.

- Radio streaming, camera vision and wake-word detection exist in earlier code

but are switched off in this build. They are not finished and are not part of

this project.


Working prototype, in daily use.


**Licence:** CC BY-SA 4.0 — build it, change it, sell what you make; credit

RoboLynk and share your changes under the same terms.


Built by RoboLynk LLC, Los Angeles.