STM32 device management: secure OTA, remote control and health monitor

You've got Mongoose running on your board, the device is on the network, and it does its job. Then you ship a few hundred of them, and a new set of questions shows up. How do I push a fix to all of them? How do I poke one device that behaves oddly? Which ones crashed last night, and where in the code?

In this article I take an STM32 Nucleo-H723ZG and hook it up to mDash, our device management service. We go through three things: firmware OTA (plain and digitally signed), remote control through custom RPC functions, and the health monitor that catches crashes and gives you a backtrace. My firmware is bare metal, no RTOS, no lwIP, no mbedTLS. But nothing here is tied to that setup: the same steps work on any MCU that Mongoose supports, bare metal or under an RTOS like FreeRTOS, ThreadX or Zephyr.

Mongoose on the device, mDash in the cloud

Quick recap if you're new here. Mongoose is a networking library that ships as two files, mongoose.c and mongoose.h. Drop them into your firmware and you get drivers, a TCP/IP stack, TLS 1.3, HTTP client and server, MQTT, Modbus, SNMP, SMTP, OTA and more. No need to glue lwIP, mbedTLS, a JSON parser and a WebSocket library together. It runs on millions of devices, from manufacturing equipment and cars to the International Space Station.

That covers the device. The fleet is a different story. Whatever your product cloud looks like, and it's usually very product-specific, you end up building the same infrastructure every time:

mDash does that part. It has been serving customers since 2017, and the idea is to run it next to your own product cloud, not instead of it. I call it a two-cloud setup, and it has some nice side effects:

The starting point

I use the nucleo-h723zg/minimal tutorial. It's a plain Makefile project with CMSIS headers and a small handwritten HAL that sets up the LED pins, the Ethernet pins and a UART for logs. The CubeMX version works the same way, and so does any other MCU that Mongoose supports, ST or not. If you have a custom board, the integration guide gets you to the same starting point.

My setup: the Nucleo's Ethernet port goes to a USB-Ethernet dongle on my workstation, and the USB cable goes to the on-board ST-Link. USART3 is wired to the ST-Link, so the serial monitor on the workstation shows the device log. I turned on internet sharing from Wi-Fi to the dongle, so the workstation runs a DHCP server there and routes the board's traffic out to the internet.

Start the serial monitor, then build and flash:

git clone https://github.com/cesanta/mongoose
cd mongoose/tutorials/stm32/nucleo-h723zg/minimal
make flash

This fetches the CMSIS headers for ARM and H7, builds and flashes. The log shows the board getting an IP address. Open it in a browser and you get "Hi from Mongoose".

main.c initialises the hardware, creates the Mongoose event manager, starts an HTTP listener, and runs a super loop with two tasks: network and blinky. The HTTP handler is small. /api/tick returns the tick counter, /api/kill crashes the device on purpose (we'll need that later), and everything else gets the greeting.

Step 1: Connect the device to mDash

Log in to mdash.net and create a device. Click the lock icon on the device card to copy its password. Now open mongoose_config.h. The minimal project already has the mDash lines in it, you only need to fill them in:

#define MG_ENABLE_MDASH 1
#define MG_MDASH_KEY "PASTE_DEVICE_PASSWORD_HERE"
#define MG_FIRMWARE_VERSION "h723.1.0.0"

TLS is already on (MG_TLS_BUILTIN), which the mDash connection needs. Run make flash again. The log now has more going on: the device connects to mDash over a secure WebSocket. In the mDash UI the device card turns green and shows the firmware version, the reset reason (unknown on the first boot) and a connectivity graph for the last 24 hours.

There's no extra code in main.c for this. mg_mgr_init() and mg_mgr_poll() take care of the mDash connection when it's enabled.

Step 2: Firmware OTA

Let's make a visible change. In blink_task() I change the blink interval from 500 to 100 milliseconds, and bump the version:

#define MG_FIRMWARE_VERSION "h723.1.0.1"

This time don't flash, just build:

make build

In mDash, click the download icon on the device card and pick firmware.bin. The device log shows it receiving data, the UI shows a progress bar, and after a reboot the card says 1.0.1. The LED blinks a lot faster.

You can do the same from a script. Get an API key at mdash.net/#/keys and:

curl -su :MDASH_API_KEY -F file=@firmware.bin https://mdash.net/api/v2/devices/DEVICE_ID/ota

The code icon on the device card copies a ready curl command for your device, so you don't have to type the URL by hand.

Step 3: Signed firmware

Plain OTA has one weak spot. If someone gets into your mDash account, they can push any firmware they like to your whole fleet. Signed updates close that.

Mongoose has a small Node script for this, resources/sign.js. No npm packages needed. Generate a key pair first:

node ../../../../resources/sign.js keygen

It writes private.pem and prints a #define with the public key. Paste that into mongoose_config.h:

#define MG_OTA_PUBLIC_KEY { \
  0x.., 0x.., ... \
}

The usual public key crypto rules apply. The public key goes into every firmware you ship, and nobody cares if it leaks. The private key signs your firmware, so keep it somewhere safe and never commit it.

Now build a signed image:

make firmware.signed.bin

This runs sign.js sign firmware.bin. The output is the same firmware with a 64-byte ECDSA P-256 signature and a 4-byte magic MGSG appended to the end, so Mongoose can tell a signed image from an unsigned one.

Push firmware.signed.bin with curl as before. Once that firmware is running, the device has the public key baked in and checks every update against it. Try pushing junk, like the Makefile:

curl -su :MDASH_API_KEY -F file=@Makefile https://mdash.net/api/v2/devices/DEVICE_ID/ota

The log says Unsigned firmware rejected. Same thing if you push a real but unsigned firmware.bin. From now on only images signed with your private key get in. Someone with your mDash password but without private.pem can't touch the firmware.

Step 4: Remote control with JSON-RPC

The device and mDash talk over a TLS-secured WebSocket, and what goes over it is JSON-RPC. JSON-RPC is about as simple as a protocol gets: a request is a JSON string with an ID, a method name and parameters, and the response is a JSON string with the same ID and a result or an error. You can watch it in the device log. Right after connecting, mDash calls Sys.GetInfo and the device answers with its firmware version, uptime, reboot reason and so on.

The fun part is that mDash also exposes a REST bridge for every device:

https://mdash.net/api/v2/devices/DEVICE_ID/rpc/METHOD

Hit that URL and mDash builds a JSON-RPC frame with that method name, puts the POST body in as parameters, sends it to the device, waits for the answer and returns it as the HTTP response. So every RPC function you add on the device is automatically an HTTP API in the cloud.

Here's a function that reads and writes any GPIO pin. It comes straight from the mDash guide. Add it to main.c:

static void mg_rpc_sys_pin(struct mg_rpc_req *r) {
  int pin = (int) mg_json_get_long(r->frame, "$.params.pin", -1);
  int val = (int) mg_json_get_long(r->frame, "$.params.val", -1);
  if (pin >= 0 && val >= 0) {
    hal_gpio_write(pin, val);  // Both pin and val are set: write pin
    mg_rpc_ok(r, "true");
  } else if (pin >= 0) {
    mg_rpc_ok(r, "%d", hal_gpio_read(pin));  // Only pin is set: read pin
  } else {
    mg_rpc_err(r, 500, "%m", MG_ESC("set pin and val"));
  }
}

And register it right after mg_mgr_init():

mg_rpc_add(&mgr.rpcs, mg_str("Sys.Pin"), mg_rpc_sys_pin, NULL);

Since the device only takes signed firmware now, build with make firmware.signed.bin and push that. The signature check passes and the device reboots into the new version.

Side note: by this point my device card had a "flapping" badge. That's mDash noticing a lot of restarts in a short time, which is exactly what you want to know about a device in the field. Here it was just me reflashing it over and over.

First, see what the device exposes:

curl -su :MDASH_API_KEY https://mdash.net/api/v2/devices/DEVICE_ID/rpc/RPC.List

You get RPC.List, Sys.GetInfo (called on every reconnect), the three OTA functions OTA.Begin, OTA.Write, OTA.End, and our Sys.Pin. Call it without arguments and you get the error from our code, "set pin and val". Now with arguments.

The pin number comes from the HAL's PIN() macro, which is (bank - 'A') << 8 | num. The board has three LEDs, on B0, E1 and B14. B0 is bank 1, pin 0, so the number is (1 << 8) + 0 = 256:

curl -su :MDASH_API_KEY https://mdash.net/api/v2/devices/DEVICE_ID/rpc/Sys.Pin -d '{"pin": 256, "val": 1}'

The green LED lights up. That's remote control in about fifteen lines of C, and you can add as many functions like this as you need.

Step 5: Health monitor

Right now the health monitor works on STM32H7 and STM32H5. The minimal project has it enabled already: the linker script has a .mg_health section, mongoose_config.h has MG_ENABLE_HEALTH 1, and main() starts with MG_HEALTH_INIT(). All that's left is to crash something.

curl http://DEVICE_IP/api/kill

That handler executes an undefined instruction. The board faults, reboots and reconnects so fast you barely notice. But the device card now says reset reason fault, shows a "crashed" badge and has a crash core download. The health monitor docs have the GDB command to read it:

arm-none-eabi-gdb -q -nx --batch \
  -ex 'set auto-load off' \
  -ex 'set debuginfod enabled off' \
  -ex 'set print frame-arguments none' \
  -ex 'set backtrace past-main on' \
  -ex 'file firmware.elf' \
  -ex 'core-file crash.core' \
  -ex 'bt'

The backtrace points right at the udf line in the /api/kill handler in main.c.

Even better, upload firmware.elf to mDash and it does the backtrace for you, right in the UI. This is my favourite bit of the whole thing. Most of the time people have no idea their devices crash in the field, let alone where. There are similar services, Memfault for example, or ESP Insights for ESP32, but here it's a config flag and a linker script line.

How it works under the hood

Open src/health.h in the Mongoose source. There's a small struct:

struct mg_health {
  char magic[4];          // MG_HEALTH_MAGIC when the record holds valid data
  uint32_t reset_reason;  // enum mg_health_reason that started this boot
  uint32_t cpuid;         // ARM CPU ID
  uint32_t data[MG_HEALTH_DATA_SIZE];
};

src/health.c defines one global instance of it in a special RAM region, and points the HardFault, MemManage, BusFault and UsageFault handlers to the same code. That code grabs three registers, SP, LR and PC. Those are all you need for a backtrace, so it ignores the rest. They go into the first three slots of data[], and the rest of the array gets a raw copy of the stack. Then it saves the CPU ID, sets the reset reason to fault and reboots the MCU.

Now, a normal global variable wouldn't survive that. Uninitialised globals live in .bss, and the startup code zeroes .bss on every boot. That's why they start at zero, and it's exactly what we don't want here. So the linker script puts the record in its own section, in a different RAM:

  /* Retained health record. Startup code must neither copy nor clear it. */
  .mg_health (NOLOAD) : { KEEP(*(.mg_health .mg_health*)) } > ram_d3

On the H723, .data and .bss live in RAM_D1. RAM_D3 is a separate block that the startup code doesn't touch, and its contents survive a software reset. After the fault handler reboots the chip, Mongoose connects to mDash, sees a valid record with the fault reason set, and sends it over. Click the cog on the device card and you can see that data: CPU ID, the three registers, then the stack. From that, mDash builds an ELF core file, which is what you download and feed to GDB.

Why upload the ELF

Picture this. A device has been running in the field for two months, and it crashes. Your repo has moved on since that release, and you didn't keep the ELF. Rebuilding the exact same ELF that matches what's on the device can be anywhere from annoying to impossible.

If you upload the ELF to mDash at release time, none of that matters. You don't need your workstation or your build environment to get a backtrace. mDash matches ELF files to devices by firmware version, so one upload covers every device running that build. That's also why you should bump MG_FIRMWARE_VERSION every time you change the code.

Wrapping up

So that's the whole loop: a device that connects to mDash, takes only firmware signed with your key, exposes any C function you like as an HTTP API, and tells you where it crashed. On the firmware side it's a few defines in mongoose_config.h, one section in the linker script, and a couple of lines in main.c.

Links: