๐Ÿค– Muse Gadget Science Fair Blueprints

Three science-fair projects rebuilt around Muse's open-source ESP32 gadgets โ€” where the AI is the variable being tested, never the thing doing the thinking. Built for a state-level fair run.

At a glance

ProjectExperiment in one lineBudgetBest for
Talking Air BuddyDoes a spoken morning CO2 briefing change window-cracking behavior more than a silent display?~$75โ€“100Strongest experiment design
The Reminder That Actually WorksWhich reminder โ€” beep, voice, or screen note โ€” gets a pill taken on time?~$40โ€“50Cheapest & fastest data
Tell It What To GrabIs voice control faster / more accurate than a joystick for pick-and-place?~$115โ€“160Best booth demo

All three keep the original workbook as a fallback โ€” if the gadget bricks, the experiment still runs on plain hardware.

Talking Air Buddy

Bedroom CO2 experiment ยท ~$75โ€“100

One-line pitch: A bedroom air-quality node you can *talk to*. Ask "How was my air last night?" and it answers from real sensor data โ€” then the experiment tests whether a *spoken* morning alert changes behavior more than a silent display.

Why this swap works

The existing co2 project is already a strong fair-test experiment (cracked vs. closed window, 14 nights, CO2 + morning grogginess). The Muse gadget replaces the plain sensor node with a voice-interactive one, and โ€” critically โ€” it adds a new testable variable: voice alert vs. silent display. The AI isn't decoration; it's the independent variable in Phase 2.

Hardware

PartWhat it doesApprox. cost (Oct 2026, US)
Waveshare ESP32-S3-Touch-AMOLED-1.75CRound touch AMOLED + mic + speaker + battery; runs the Muse Device SDK~$30โ€“50
SCD40 CO2/temp/humidity breakout (genuine Sensirion)The actual sensor (400โ€“2000 ppm range, ยฑ(40 ppm + 5%))~$30โ€“45 generic; ~$45โ€“90 SparkFun/Adafruit
USB-C cable, decent gauge (20AWG or thicker)Overnight power โ€” thin cables brownout the ESP32-S3 during Wi-Fi bursts~$8
Jumper wires + small stand/enclosureMount the sensor away from the board's heat~$10
**Total****~$75โ€“100**

Buyer beware: fake SCD40 chips are common. Verify genuineness before trusting data: run it outdoors 20 minutes โ€” a real one settles near 400โ€“420 ppm. A clone reads high/low or never stabilizes. Buy from a reputable seller.

Software setup

  • 1. Get an SDK token at gadgets.muse.ai and clone Meta's ESP32 Device SDK (Apache 2.0, on GitHub).
  • 2. Flash the gadget firmware to the Waveshare board (follow the SDK's board-specific guide).
  • 3. Wire the SCD40 over I2C: 3.3V โ†’ VDD, GND โ†’ GND, GPIO SDA/SCL per the board's pinout.
  • 4. Configure the gadget's persona: "You are a bedroom air-quality buddy. Answer questions about last night's CO2 readings in plain kid-friendly language. Suggest cracking the window when readings were high."
  • 5. Keep Wi-Fi on overnight โ€” the voice brain lives in the cloud.

The experiment (two phases)

Phase 1 โ€” the original fair test (reuse the `co2` workbook)

  • Question: Does cracking the bedroom window keep overnight CO2 near outdoor levels?
  • Independent variable: window cracked vs. closed, alternating nights, 14 nights total.
  • Dependent variables: overnight average + peak CO2 (ppm), morning grogginess (1โ€“5).
  • Controls: same room/sleeper, same sensor spot and height, same bedtime window, fan/heater setting fixed, minute-by-minute logging.
  • Gadget role: live CO2 gauge on the round display + voice Q&A over the data ("What was my peak last night?").

Phase 2 โ€” the gadget's own experiment (the new science)

  • Question: Does a *spoken* morning air report make me crack the window more often than a silent display?
  • Hypothesis (If/Then/Because): If the gadget speaks a morning CO2 briefing, then I will crack the window on more nights and my average overnight CO2 will be lower than with a silent display, because a spoken alert is harder to ignore than a number on a screen.
  • Independent variable: alert mode โ€” spoken briefing vs. silent display, alternating in 2-week blocks (ABBA order: silent, voice, voice, silent โ€” or reversed).
  • Dependent variables: % of nights the window was actually cracked (compliance), average overnight CO2 per block.
  • Controls: same as Phase 1, plus: briefing wording kept identical every morning, briefing plays within 5 minutes of wake-up, grogginess rating recorded before hearing the briefing.
  • Data table: one row per night โ€” date, mode, window (cracked/closed), avg CO2, peak CO2, grogginess, notes.

Analysis

  • Phase 1: bar chart (cracked vs. closed avg CO2) + scatter plot (CO2 vs. grogginess) โ€” same as the existing workbook.
  • Phase 2: bar chart (voice-block vs. silent-block compliance % and avg CO2). State honestly: 14 nights per block is small โ€” report the difference and the range, not certainty.
  • Conclusion sentence starter: "My hypothesis was ___ because the voice-alert block averaged ___ ppm over ___% cracked nights vs. ___ ppm over ___% cracked nights in the silent block."

Judge Q&A (new, gadget-specific)

  • "Isn't the AI doing the thinking for you?" โ†’ "No โ€” the gadget only reports and reminds. I designed the experiment, chose the variables, and analyzed the data. The voice alert is the thing being tested, not the thing doing the testing."
  • "Why alternate the alert modes in blocks instead of nightly?" โ†’ "A nightly switch would confuse the habit I'm measuring. Two-week blocks let a routine form, and the ABBA order balances out weather changes."
  • "What if the sensor is wrong?" โ†’ "I verified it outdoors at ~420 ppm before starting, and I report its ยฑ(40 ppm + 5%) accuracy on the board."
  • "Could you have gotten the same result with a phone alarm?" โ†’ "That's exactly why the silent-display block is the control โ€” it isolates whether the *voice* matters, not just the reminder."

Booth demo (the judge magnet)

A judge presses the button and asks, "Was last night stuffy?" The gadget answers from real logged data and shows the gauge on its round display. Live, personal, and every word is backed by the data table on the board behind it.

Risks and honest limits

  • Hacker-grade hardware. Meta's own page warns about bricked boards. Budget a spare ~$30 and a debugging weekend. This is a build skill the judges will respect โ€” if it survives.
  • Wi-Fi dependency. No Wi-Fi, no voice. The sensor logging must be local so a dead network never loses a night's data.
  • SCD40 auto-calibration assumes regular fresh-air exposure; a sealed bedroom for weeks can drift it. The weekly outdoor check doubles as calibration sanity.
  • Small samples. 14 nights per condition can only show big effects. Say so on the board โ€” honesty about limits scores points.
  • Phase 2 measures the kid's own behavior โ€” that's fine, but disclose it: "n=1, myself." Suggest the follow-up: run it on a sibling or friend.

Timeline

  • Weekend 1: order parts, flash firmware, wire sensor, verify ~420 ppm outdoors.
  • Weekend 2: persona config, voice Q&A testing, booth-demo dry run.
  • Weeks 1โ€“2: Phase 1 data (14 nights).
  • Weeks 3โ€“6: Phase 2 data (two 2-week blocks).
  • Final weekend: charts, board, Q&A practice.

What stays from the existing `co2` workbook

Everything in Phase 1 โ€” hypothesis_fair, variables, control, data table, observations, analysis steps, board layout โ€” carries over unchanged. This blueprint only *adds* the gadget layer and Phase 2. The original workbook remains the fallback: if the gadget bricks, the experiment still runs on the plain sensor node.

The Reminder That Actually Works

Smart pill reminder ยท ~$40โ€“50

One-line pitch: A pocket-size gadget tests which reminder gets a pill taken on time โ€” a plain beep, a silent screen note, or a voice that talks to you. The AI isn't the assistant; it's the thing being tested.

Why this swap works

The existing meds project already asks the right question (do reminders improve on-time doses?). The Muse gadget makes it a *three-way* comparison โ€” beep vs. voice vs. visual note โ€” which is a richer, more judge-worthy experiment than the original's reminder-vs-nothing design. Human-factors experiments like this are rare at school fairs and read as genuinely useful, not just technical.

Important ethics note: the volunteer must be a consenting adult (a parent or grandparent on a real daily medication). The kid designs and runs the experiment but is never the subject. Never put the volunteer's name or the medication name on the board, in the log, or in the video โ€” use "Volunteer A" and "daily morning medication."

Hardware

PartWhat it doesApprox. cost (Oct 2026, US)
M5Stack StickS3Pocket-size ESP32-S3 gadget: color screen, mic, speaker, battery, officially supported by the Muse Device SDK~$22
7-day pill organizer (large compartments)The physical pill box; the gadget sits beside it~$8
Magnetic reed switch + small magnet (optional)Detects when the pill-box lid opens = dose-taken timestamp, no human logging needed~$6
USB-C cableCharging~$8
**Total****~$40โ€“50**

Why the StickS3: it's the cheapest officially supported Muse gadget (~$22), has everything needed (mic + speaker + screen + battery), and its tiny size means it can live on a nightstand next to the pill box. No e-ink needed โ€” the color screen shows the note fine.

Software setup

  • 1. SDK token from gadgets.muse.ai; flash the Muse Device SDK firmware to the StickS3.
  • 2. Configure three reminder "personalities":
  • Beep mode: plays a simple tone at dose time, no words, no screen text beyond a pill icon.
  • Voice mode: speaks a short briefing โ€” "Good morning. It's 8 AM โ€” time for your morning pill." Kid-configurable wording, kept identical every day.
  • Note mode: silent; screen shows "8:00 AM โ€” take your pill" until the lid opens.
  • 3. If using the reed switch: wire it to a GPIO; log lid-open timestamps to the gadget's storage (local logging โ€” a dead Wi-Fi night must never lose data).

The experiment

  • Question: Which reminder type produces the highest on-time dose rate?
  • Hypothesis (If/Then/Because): If the reminder is a spoken voice message, then on-time doses will be higher than with a beep or a silent screen note, because a voice naming the action is harder to dismiss than a tone you can tune out or a note you can walk past.
  • Independent variable: reminder type โ€” passive (no reminder, just logging), beep, voice, silent note. Four 1-week blocks in ABBA-ish order (e.g., passive, beep, note, voice โ€” then repeat in reverse for a second volunteer or a second month).
  • Dependent variable: on-time adherence โ€” % of doses taken within 1 hour of schedule.
  • Controls: same volunteer, same medication schedule, same dose times, same organizer layout, same 4-week window (no travel/holidays), reminder wording fixed, reminder fires at the exact dose time every day.
  • Data table: one row per dose โ€” date, block type, scheduled time, taken time, on-time (yes/no), notes (e.g., "volunteer was in the shower").

Analysis

  • Bar chart: on-time % by reminder type (4 bars). This is the money chart.
  • Honest stats: with ~7 doses per block, differences need to be large to mean anything. Report the percentages and say plainly what the sample can and can't prove.
  • Conclusion starter: "My hypothesis was ___ because the voice block reached ___% on-time vs. ___% for beep, ___% for the note, and ___% passive."

Judge Q&A

  • "Isn't this just testing an alarm clock?" โ†’ "A beep IS an alarm clock โ€” that's the point of including it as a comparison. The experiment asks whether a *voice* beats the alarm clock, and by how much."
  • "One volunteer isn't enough." โ†’ "Agreed โ€” and I say that on the board. This is a pilot. The next step is 5 volunteers, and my protocol is written so anyone can repeat it."
  • "How do you know the dose was actually taken, not just the lid opened?" โ†’ "I don't, fully โ€” lid-open is a proxy. I disclose that as a limit. A camera would verify but would be invasive; I chose the less invasive measure and named the trade-off."
  • "Did the AI do the experiment for you?" โ†’ "The gadget delivers the reminders. I designed the comparison, ran the 4 weeks, and analyzed the data. The AI is the treatment, not the researcher."

Booth demo

The judge presses the StickS3 button and hears the actual voice reminder, then sees the week's adherence chart on the board. Bonus: a second StickS3 in beep mode for the judge to compare โ€” "which one would get YOU to take the pill?"

Risks and honest limits

  • Real medication = real responsibility. If the volunteer misses doses, that's their health โ€” the experiment must never pressure or shame. The passive week is ethically fine (it's their normal routine), but get explicit adult consent in writing.
  • n=1. One volunteer, four weeks. Frame it as a pilot study, always.
  • Hawthorne effect: the volunteer knows they're being measured, which alone can improve adherence. The passive block helps, but name the effect on the board โ€” judges love seeing it named.
  • Hacker-grade hardware (same bricking caveat as the CO2 build). The reed-switch logging is the fiddliest part โ€” test it for a full week before the experiment starts.
  • Privacy: no names, no medication names, no photos of the volunteer. Ever.

Timeline

  • Weekend 1: order parts, flash StickS3, configure the three reminder modes.
  • Weekend 2: reed-switch wiring (if used), 3-day dry run of logging.
  • Weeks 1โ€“4: the four experiment blocks.
  • Final weekend: charts, board, Q&A practice.

What stays from the existing `meds` workbook

The core protocol โ€” on-time adherence definition, 1-hour window, control list โ€” carries over. This blueprint upgrades the two-condition design (reminder vs. nothing) to a four-condition comparison and swaps the generic reminder for the Muse gadget's voice.

Tell It What To Grab

Voice-controlled robot arm ยท ~$115โ€“160

One-line pitch: A robot arm that obeys spoken commands โ€” "pick up the red block, put it in the blue zone." The experiment: is talking to a robot faster and more accurate than driving it with a joystick?

Why this swap works

The existing arm project already compares two control methods (motion glove vs. joystick). Swapping the glove for *voice* makes it a cleaner, more modern comparison โ€” and voice control of physical hardware is exactly the kind of demo that stops judges mid-aisle. The experiment stays a fair human-factors test: same arm, same task, only the control method changes.

Hardware

PartWhat it doesApprox. cost (Oct 2026, US)
4โ€“6 DOF servo robot arm kit (MG996R-class servos, aluminum or acrylic frame)The arm itself; many kits include the servo controller board~$60โ€“100
M5Stack StickS3The voice puck: mic + speaker + screen, officially supported by the Muse Device SDK; listens for commands and relays them to the arm~$22
ESP32 dev board (plain, e.g. ESP32-WROOM) โ€” if the arm kit has no smart controllerReceives commands from the StickS3 over Wi-Fi and drives the servos~$8โ€“12
Joystick module (KY-023 style)The comparison control method~$5
Colored wooden blocks + two marked zones (red/green/blue)The pick-and-place task props~$10
5V/2A+ power supply for servosServos brown out on weak USB power โ€” dedicated supply, always~$10
**Total****~$115โ€“160**

Cheapest path: a ~$60 4-DOF acrylic arm kit + StickS3 + joystick. The ~$100โ€“160 end buys a 6-DOF aluminum kit with smoother motion (fewer "the arm physically couldn't reach it" excuses in the data).

Software setup

  • 1. Assemble the arm kit per its manual; verify every joint moves through its full range before writing any code.
  • 2. Flash the Muse Device SDK to the StickS3 (SDK token from gadgets.muse.ai). Configure a constrained command vocabulary โ€” the gadget should map phrases like "pick up the red block," "move left," "open gripper," "put it in the blue zone" to fixed arm commands. Constrain it deliberately: an open-ended chatbot controlling servos is a safety and fairness problem.
  • 3. The StickS3 sends commands to the arm's ESP32 over local Wi-Fi (simple HTTP or MQTT on the home network โ€” no cloud round-trip for the motion itself, so latency stays low and consistent).
  • 4. Joystick mode: the same ESP32 reads the KY-023 and drives the same servo functions โ€” identical motion code, only the input differs. This is what makes the comparison fair.

The experiment

  • Question: Is voice control faster and more accurate than joystick control for a pick-and-place task?
  • Hypothesis (If/Then/Because): If I control the arm by voice, then pick-and-place rounds will be *slower* than with the joystick but will have *fewer errors*, because speaking a command ("put it in the blue zone") aims the whole task at once while a joystick needs constant fine steering โ€” OR the reverse; state your honest guess before running.
  • Independent variable: control method โ€” voice vs. joystick.
  • Dependent variables: time in seconds to move 5 blocks to their target zones; error count (dropped or misplaced blocks); a 3-question operator survey (which felt easier? which felt more tiring? which would you trust with something fragile?).
  • Controls: same arm, same 5 blocks, same start/target positions every round; same operator for all rounds; 10 rounds per method (20 total), alternating voice/joystick to spread practice effects evenly; blocks reset to identical spots before each round; command vocabulary frozen before trials (no adding phrases mid-experiment); Wi-Fi latency checked before each session.
  • Data table: one row per round โ€” round #, method, time (s), errors, notes ("voice misheard 'blue' as 'glue'", "joystick drift on joint 3").

Analysis

  • Bar chart: average time by method. Bar chart: average errors by method. Side by side โ€” the trade-off is the story.
  • Scatter plot (optional, judge candy): each dot is a round, x = round number, y = time, colored by method โ€” shows the learning curve flattening for both.
  • Survey results as a simple 3-row table.
  • Conclusion starter: "My hypothesis was ___ because voice averaged ___ seconds with ___ errors per round vs. joystick's ___ seconds with ___ errors."

Judge Q&A

  • "Isn't the AI doing the work?" โ†’ "The AI translates my words into arm commands โ€” but I designed the comparison, ran all 20 rounds, and measured everything. The voice interface is what's being tested."
  • "What if it mishears you?" โ†’ "That's data, not failure โ€” mishears are logged as errors under the voice column. A real product would have this exact problem, which is why the experiment matters."
  • "Why only one operator?" โ†’ "One operator keeps skill level constant, which is the fair-test requirement. The follow-up is three operators to check the result generalizes."
  • "Could the joystick operator just be better at joysticks?" โ†’ "Alternating the rounds spreads practice evenly, and the learning-curve plot shows both methods improving โ€” if anything, that makes the comparison conservative."

Booth demo (the showstopper)

This is the best demo of the three gadget projects. A judge says "move the green block to the red zone" โ€” the StickS3 confirms out loud, the arm does it. Then the judge tries the joystick themselves and immediately feels the difference. Hands-on judging beats poster-reading every time.

Safety for the booth: cap servo speed in the demo firmware, keep the arm's reach inside a marked boundary, and have a big red stop button (the StickS3's front button = all-stop). Mention the safety design on the board โ€” judges notice.

Risks and honest limits

  • The arm kit is the long pole. Assembly + calibration is a full weekend before any AI work starts. Buy the kit first; don't order the StickS3 until the arm moves.
  • Voice latency and mishears are the experiment's noise โ€” log them honestly rather than re-running "bad" rounds. Cherry-picking rounds is the fastest way to lose a judge's trust.
  • n=1 operator. Same disclosure as the meds project: pilot study, protocol written for repeatability.
  • Hacker-grade SDK (same bricking caveat). Keep the joystick firmware as the fallback โ€” if the voice stack fails the week before the fair, the original glove/joystick comparison still runs.
  • Servo power: underpowered servos jitter and stall, which corrupts timing data. The dedicated 5V supply isn't optional.

Timeline

  • Weekend 1: order arm kit; assemble and calibrate (no AI yet).
  • Weekend 2: joystick control working; baseline timing runs to shake out the task design.
  • Weekend 3: StickS3 voice setup; constrained vocabulary testing.
  • Week 4: the 20 trial rounds (spread over several days to avoid fatigue effects).
  • Final weekend: charts, board, Q&A practice, booth safety check.

What stays from the existing `arm` workbook

The task design (5 blocks, timed rounds, error counting, alternating methods, 10 rounds per method) carries over almost unchanged โ€” only the glove becomes voice. If the voice stack proves unreliable, the workbook's original glove-vs-joystick version is the intact fallback.