Charter 21 · system · reserve project

Spine Doctor

Distributed Health Monitor

Watch the watchers—and learn why silence becomes evidence only after someone promises when to speak.

Difficulty
★★★
Prefix
sd
Zone
workshop
Type
spine-doctor
Device
sd-01
Build time
10 hours

Mission

Build this

Build a physical clinic that tracks expected periodic evidence from selected nodes, separates late from silent, and raises one calm flag for a documented outage.

One kit, one team repository.

Borrowed parts are labelled below and have a fallback. The registered device ID is fixed; sensing and behaviour decisions remain yours.

Bill of materials

Parts

SourcePartFallback
● kitOLED—
● kitRGB LED and 220 Ω resistors ×3—
● kitActive buzzer module—
● kitSG90 servo—

Disconnect USB before rewiring. Motors, the relay, and the servo need appropriate power and a shared ground. Never drive an actuator from the ESP32 3V3 pin.

Backup assignment · your invitation

Become useful when everyone else’s project becomes mysterious

Spine Doctor watches a small, declared set of garden devices and asks whether their expected evidence is still arriving. It shows who is timely, who is late, and who has crossed a documented silence limit.

The danger is false confidence. A quiet status topic may belong to a healthy device whose state has not changed. A chatty device may be online while publishing nonsense. Your monitor must say exactly which kind of health it observes.

Clinical humility

“No expected message arrived” is an observation. “The sensor is broken” is one possible diagnosis.

The central idea

Online, fresh, and correct are three different claims

connection

Can it reach the garden?

Broker and backend evidence may show recent communication.

freshness

Did expected data arrive?

A periodic topic can be compared with its documented interval.

correctness

Does the value make sense?

This needs project-specific validation that Spine Doctor usually does not possess.

Your required claim is freshness for selected periodic topics. Status messages provide useful context, but because many publish only when state changes, silence on status alone is not evidence of failure.

Write the care plan

Every monitored device deserves its own clock

DevicePeriodic topicExpected intervalGraceLate afterWhy this topic?
gm-01…/temperature60 s___ s___ sregular climate reading
cc-01…/temperature5 min___ min___ minregular compost reading
team chooses…____________
Do not monitor an on-change topic as though it were periodic.

A counter may remain unchanged for hours. A warning status may remain ok for days. Choose a periodic measurement or make an explicit heartbeat agreement with that team.

Wiring

Build a calm clinic, not a permanent alarm

Show monitored devices and message ages on the OLED before adding the flag. Use colour and one brief sound only when the overall state changes; repeated beeping does not improve diagnosis.

The power rule: The OLED, RGB logic, and verified buzzer module use 3V3. Each LED colour needs 220 Ω. The servo uses a separate 5 V supply. Join all grounds.

Bench referenceOpen the exact wiring map
WiringMessage-age display, three-state clinic light, brief cue, and system flag
ESP323V3OLED (SSD1306)VCCGNDOLED (SSD1306)GNDGPIO 21OLED (SSD1306)SDAGPIO 22OLED (SSD1306)SCLGPIO 25RGB LEDR via 220 ΩGPIO 26RGB LEDG via 220 ΩGPIO 27RGB LEDB via 220 ΩGNDRGB LEDcommon cathodeGPIO 33Active buzzer moduleSIG3V3Active buzzer moduleVCCGNDActive buzzer moduleGNDGPIO 18SG90 servosignal (orange)external 5 VSG90 servopower (red)GNDSG90 servoground (brown)
ESP32 pinPartPart markingCarries
3V3OLED (SSD1306)VCC3V3 power
GNDOLED (SSD1306)GNDGround
GPIO 21OLED (SSD1306)SDAI²C bus
GPIO 22OLED (SSD1306)SCLI²C bus
GPIO 25RGB LEDR via 220 ΩDigital in/out — common-cathode LED; invert for common-anode
GPIO 26RGB LEDG via 220 ΩDigital in/out
GPIO 27RGB LEDB via 220 ΩDigital in/out
GNDRGB LEDcommon cathodeGround
GPIO 33Active buzzer moduleSIGDigital in/out — brief transition cue only
3V3Active buzzer moduleVCC3V3 power
GNDActive buzzer moduleGNDGround
GPIO 18SG90 servosignal (orange)PWM to actuator — system flag
external 5 VSG90 servopower (red)5V power — separate supply
GNDSG90 servoground (brown)Ground — join servo and ESP32 grounds

Never power the servo from ESP32 3V3. Confirm the buzzer type and current before wiring. A health monitor must remain quiet enough for teams to leave it running.

  • 3V3 power
  • Ground
  • I²C bus
  • Digital in/out
  • PWM to actuator
  • 5V power

The map assumes a common-cathode RGB LED and small three-pin active-buzzer module. Invert RGB logic for common-anode hardware.

Give silence a fair trial

Late should come before silent

never seen

No baseline exists. Do not count the device alive or failed.

watching

Expected messages arrive inside the documented interval plus grace.

late

The expected time has passed. Show uncertainty before raising the flag.

silent

A longer evidence threshold is crossed. Raise the flag once.

recovered

A new valid message arrives. Record recovery delay and lower the flag.

Your team decides:

  • How much grace reflects real network variation rather than wishful waiting?
  • Does one silent node make the overall status warning or error?
  • How are “never seen” and “was seen, now silent” displayed differently?
  • What evidence is required before a recovered node rejoins the alive count?

Your controlled outage ward

Break one known connection at a time

  1. Collect a healthy baseline.

    Run the selected devices through enough expected intervals to measure ordinary message delay and occasional gaps.

  2. Schedule short and long outages.

    With the owning team present, disconnect one test device for durations below and above its thresholds.

  3. Record the complete response.

    Measure late detection, silent detection, false alarms on other nodes, and recovery after the first valid message.

  4. Repeat after one rule change.

    Use the same outage script so lower false alarms cannot hide a much slower diagnosis.

DevicePlanned outageLate delaySilent delayRecoveryOther false alarms
gm-01___ min___ s___ s___ s___
cc-01___ min___ s___ s___ s___
short outagebelow threshold___no______

When the doctor misdiagnoses the garden

Disagreement tells you which health model you built

A healthy event counter is declared dead

The monitored topic publishes only when something happens. Replace it with a periodic topic or negotiate a periodic heartbeat.

Every node fails at the same moment

The shared network, broker, or Doctor itself may be unavailable. Publish an error for the monitor’s own connection instead of blaming every patient independently.

The flag chatters during brief delays

The grace region is too narrow or has no late state. Measure normal arrival variation before changing thresholds.

The dashboard says alive while Doctor says silent

They may observe different messages or use different timeouts. Compare the exact last-seen times and definitions before calling either wrong.

A recovered node leaves the alarm raised

The overall state or alive count was not recomputed after its valid message. Derive all outputs from one current patient table.

Choose your systems question

What kind of Doctor team will you become?

The fairness designers

Build device-specific expectations and show why one universal silence timeout rewards chatty nodes and punishes slow ones.

The outage scientists

Make detection delay, recovery, and false alarms the experiment across a carefully scripted set of failures.

The model comparers

Compare Doctor, dashboard last-seen state, and API health. Explain disagreements through definitions and observation points.

A calm way through the build

Collect four small wins

  1. sd-01 says hello.

    Prove the monitor can report its own status first.

  2. One periodic topic shows a growing message age.

    Record arrival time without using a blocking timer.

  3. One short outage becomes late, not dead.

    Add grace and a patient-specific silence threshold.

  4. One long outage and recovery move the flag once each.

    Add the remaining patients, then run the fixed script.

The garden handshake

Share what you found

These names are the rigid part of the project. They let another team find your work without knowing what you called the variables in your code.

Every 60 seconds

Patients currently fresh

garden/workshop/spine-doctor/sd-01/nodes-alive

Unit: devices

When state changes

Monitor operating state

garden/workshop/spine-doctor/sd-01/status

Unit: enum

Listen beyond your own device.

garden/+/+/+/status

Wildcard status messages provide declared state context, not a universal heartbeat. Determine aliveness from selected periodic topics or explicit heartbeat agreements with documented intervals.

Finish line

Ready to introduce to the garden

  • sd-01 stays online and publishes fresh-node count every minute plus a registered monitor status.

  • The reject feed stays clear after the final code starts.

  • Each monitored node has a documented periodic evidence topic, expected interval, grace, and silence threshold; on-change status is not treated as heartbeat.

  • Display, light, brief cue, and flag derive from one patient table without blocking messages or repeating alarms.

  • The build log contains a healthy baseline, planned short and long outages, late and silent detection delays, recovery time, false alarms, and a repeated script after one change.

  • The monitor distinguishes its own network failure from patient silence, the servo is safely powered, sound is restrained, and the label can be scanned.

Spine Doctor is trustworthy when its alarm names the evidence it lost, the clock it expected, and the limits of the diagnosis—not merely the device it suspects.

Backbone now

Live status

Updates from the same public event stream

Checking sd-01…