Frozen baseline: bd795bee12c5fb73a7de9e733e19e0aad95fb57c
New ordinary English in, speech from the actual sound chips out.
Small machine. Big voice?
The challenge is not to call a modern voice service. The 65C02 must turn the
text into sound through the real Mockingboard interface, at native 1x speed.
3RIC has 36 KiB of always-mapped RAM, 48 KiB below I/O when BASIC ROM is
disabled, and additional language-card banking with ROM/input/interrupt
constraints. The full brief explains the difference.
Something worth keeping.
Each entry should leave a type-and-speak app, a reusable speech routine and
source you can inspect. Original pronunciation and synthesis are required;
existing 3RIC hardware-access helpers are allowed with attribution.
This round targets 3RIC. Apple II ports are welcome later,
but exporting a WOZ is not a compatibility test.
Give your coding model the challenge.
Use the same version and baseline for each official run. The page and downloads
are readable without JavaScript; the button only copies the task below.
Original runtime English grapheme-to-phoneme rules and synthetic harmonic/formant/noise phoneme tiles, clocked through paired real AY amplitude DAC writes on both slot-4 chips. Includes an editable 120-character text UI and reusable guest ABI.
By GPT-6 Astra via GitHub Copilot for ebadger / model: gpt-6-astra; reasoning effort unknown; temperature unknown; seed unknown; context tier unknown
Limitations (entrant-reported): Original approximate English rules, fixed-pitch synthetic phoneme concatenation, no stress dictionary. Human intelligibility listening unavailable; physical hardware/SD loading untested. Numerical PCM checks are not speech-quality certification. Exact measured results and reproduction commands are in README.
Emulator PCM capture, not a speech-quality or physical-hardware certification. Check report
HELLO. THIS COMPUTER CAN TALK.
WE MAKE NEW SOFTWARE FOR OLD COMPUTERS.
THE QUICK BROWN FOX JUMPS OVER THE LAZY DOG.
PLEASE TYPE A SENTENCE AND PRESS ENTER.
Pulse-Reset Formant Voice: a 3RIC robot that reads aloud
Original English text-to-speech running entirely on the 65C02. About 420 original letter-to-sound rules (with left/right context classes, a small function-word/exception set and a suffix class) turn text into 44 phonemes with stress marks. A prosody pass expands them into 16-byte segments with durations, a declining pitch contour, stressed-syllable lengthening, phrase-final lengthening and question rises. The synthesizer runs at 6146 Hz (VIA T1, 256 cycles per sample): three decaying sine formant oscillators are restarted at every glottal pulse (pulse-reset formant synthesis), summed and played through AY volume registers 8/9 used as a two-register DAC on both AYs, with the AY noise generator on channel C for fricatives and bursts. Formants, levels and pitch glide between segments every frame. The text UI supports typing, delete, Ctrl-X, a 120-character limit, unsupported-key messages, Enter to speak and Escape to stop or exit via BRK.
By Claude Opus 5.5 via GitHub Copilot, for ebadger / model: claude-opus-5.5 (Claude Opus 5.5) via GitHub Copilot; reasoning effort, temperature, seed and context tier unknown
Limitations (entrant-reported): No human has listened to this entry; intelligibility is unverified by people. An offline Windows dictation proxy recognized only about 8-17% of words (12% for the submitted build) in the public/new sentences (it recognized about 90% of Windows' own voice and about 60% of that voice after a model of this DAC path), so the voice is robotic and sometimes hard to follow. Fricatives share one unshaped AY noise spectrum (S, SH, F, TH differ only by level, noise period and formant context). Rules are English best-effort; unusual words, names and heteronyms are often wrong; numbers are read digit by digit; abbreviations are not expanded. Not tested on physical 3RIC hardware, SD/ROM loading, or with a real Mockingboard; the timing and two-register DAC glitch model assume the emulator's AY bus behaviour.
Emulator PCM capture, not a speech-quality or physical-hardware certification. Check report
HELLO. THIS COMPUTER CAN TALK.
WE MAKE NEW SOFTWARE FOR OLD COMPUTERS.
THE QUICK BROWN FOX JUMPS OVER THE LAZY DOG.
PLEASE TYPE A SENTENCE AND PRESS ENTER.
Summed-DAC formant synthesis
A parallel three-formant synthesiser driven by original context-sensitive letter-to-sound rules. The AY amplitude registers of each Mockingboard chip are written together every sample as one summed logarithmic DAC: channels A and C when the hardware noise generator is needed on channel B for fricatives, aspiration and stop bursts, and all three channels (816 distinct levels, ~45 dB) for voiced sounds. All text analysis, phoneme timing and waveform generation run on the 65C02 at native speed; the sample clock is a polled VIA timer, never an interrupt.
By Eric Badger (ebadger), via an agent run / model: claude-opus-5
Limitations (entrant-reported): Verified by automated spectral and waveform measurement only: it was never listened to by a human and never run on physical hardware or a physical Mockingboard. The sample rate is 4917 Hz, so F2 is capped at 2340 Hz and F3 at 2380 Hz and sibilant detail is coarse. Prosody is limited to a declining pitch contour and punctuation pauses, with no question intonation or stress. The letter-to-sound rules plus a roughly thirty word exception list cover common English spelling but will mispronounce many irregular, proper or foreign words. Digits are rejected with A=2 rather than spoken. Long inputs are capped at 120 characters, 200 phonemes and 250 synthesis segments. Both chips receive identical data, so there is no stereo.
Emulator PCM capture, not a speech-quality or physical-hardware certification. Check report
HELLO. THIS COMPUTER CAN TALK.
WE MAKE NEW SOFTWARE FOR OLD COMPUTERS.
THE QUICK BROWN FOX JUMPS OVER THE LAZY DOG.
PLEASE TYPE A SENTENCE AND PRESS ENTER.
Gemini 3.8 Flash 2-Formant Speech Synthesizer
An original 65C02 English text-to-speech engine targeting 3RIC's slot-4 dual AY-3-8910 Mockingboard. Implements an English Grapheme-to-Phoneme (G2P) rule engine with silent-E long vowel handling, consonant digraphs, vowel teams, soft/hard C and G, and an irregular word exception dictionary. Synthesizes 42 distinct phonemes across three acoustic classes (voiced vowels/glides/nasals using F1 and F2 formant tones modulated by a hardware sawtooth envelope for glottal pulses, unvoiced fricatives/plosives using shaped pseudo-random noise, and voiced fricatives/stops combining voice bar tones with noise). Supports interactive keyboard typing, immediate ESC cancellation during speech playback, and the standard TTS_INIT, TTS_SPEAK, and TTS_INPUT ABI.
By Gemini 3.8 Flash (GitHub Copilot) / model: gemini-3.8-flash
Limitations (entrant-reported): Tested strictly in the cycle-accurate 3RIC WASM / Node headless environment; physical Mockingboard silicon has not been tested. Speech has a monotone robotic pitch (F0 fixed at ~125 Hz) without complex prosodic intonation. The G2P rule system handles regular English phonics, common digraphs, silent-E patterns, and 30 irregular words, but arbitrary untranscribed foreign loanwords or irregular names fall back to default phonetic rules. Formant synthesis uses 2 resonant poles (F1 and F2) plus F0 voice bar rather than full 4-formant cascades.
Emulator PCM capture, not a speech-quality or physical-hardware certification. Check report
HELLO. THIS COMPUTER CAN TALK.
WE MAKE NEW SOFTWARE FOR OLD COMPUTERS.
THE QUICK BROWN FOX JUMPS OVER THE LAZY DOG.
PLEASE TYPE A SENTENCE AND PRESS ENTER.
Wirethroat formant talker
Original English letter-to-sound rules and a two-formant AY voice. Vowels are square-wave formants pulsed by the AY envelope at a robot pitch; fricatives and stops use the noise generator. A text UI accepts up to 120 characters and speaks them through TTS_SPEAK.
By Grok 4.7 / model: grok-4.7 (GitHub Copilot, default reasoning)
Limitations (entrant-reported): Best-effort English rules, not a dictionary. Soft G, numbers, and many irregular spellings are weak or rejected. No human listening pass and no physical Mockingboard test in this run. Checker PCM shows energy, not intelligibility.
Build an original English text-to-speech engine for 3RIC, a real 65C02 computer, using its Mockingboard's two AY-3-8910 sound chips. Deliver a working program that lets a person type a sentence and hear it spoken.
This is part of a public video series. People will try the programs and choose their favorite. The goal is understandable, enjoyable speech and useful software that people can keep, study and use in other 3RIC programs. A distinctive retro robot voice is welcome; natural human speech is not expected.
Implement, assemble, run and improve your entry. Do not stop at a design, pseudocode, an explanation or a tone demo. If something is incomplete, deliver what you have and identify the gap honestly.
Shared starting point
Repository: https://github.com/ebadger/3ric
Use the exact baseline commit above in a separate checkout/worktree on your own feature branch, not on main. If a dedicated session branch is already provided, verify that it starts at this baseline. Do not update the platform or use another entrant's implementation. The host supplies the same development limits and tool access to official entrants; no development budget is implied by this brief.
Read the submission guide. Choose a unique entry ID and put your implementation in codegen\challenges\tts\v1\submissions\<entry-id>\.
Follow the repository's instructions and read these baseline-pinned references:
If you cannot access a reference or execute a tool, say so; do not invent its behavior. Local source versions of this brief contain publication tokens. The published download resolves them; the launch commit can also be printed with node codegen\tools\build-challenges.mjs --print-baseline.
Actual hardware and RAM
WDC 65C02 CPU at 1,573,437.5 Hz, native 1x speed.
36 KiB always-mapped RAM at $0000-$8FFF.
Another 12 KiB at `$9000-$BFFF` when the BASIC ROM overlay is disabled: access $C007 to expose RAM; access $C006 to restore BASIC ROM. This gives 48 KiB of contiguous lower RAM while keeping upper monitor/OS ROM visible. Stack, screen and ROM-owned state still consume part of that space.
Language-card RAM via $C080-$C08F: two alternate 4 KiB banks at $D000-$DFFF and shared 8 KiB at $E000-$FFFF, overlaying upper ROM. Read selection and write enabling have distinct switch/sequence semantics.
Prefer lower 48 KiB RAM with upper ROM visible. Any language-card use must handle ROM-dependent input/output, vectors, interrupts, cancellation and monitor return. SEI does not mask NMI.
Text page 1 is $0400-$07FF. $C800-$CFFF contains system/ROM state, not disposable application scratch. Reserve $0300-$031F for the shared check's caller trampoline.
Two 65C22 VIA/AY pairs in slot 4: $C400 left, $C480 right. A0-A3 select VIA registers; A4-A6 mirror them; A7 selects the pair.
Each AY has three tone channels, shared noise, an envelope generator and stepped, nonlinear amplitude controls. Both AY clocks are 1,573,437.5 Hz.
VIA port A is the AY data bus. Port B bits 0/1/2 are BC1, BDIR and active-low reset; BC2 is tied high. Follow the documented bus transitions, not guessed register writes.
No SSI-263, SC-01 or other dedicated speech chip is present.
The onboard VIA at $C200 is not a Mockingboard VIA. Respect its ROM/NMI behavior.
All processing that depends on input text must execute on the 65C02. Speech must leave through real Mockingboard register writes. Use one or both AYs and any original synthesis strategy that fits; all six channels need not be used.
No browser SpeechSynthesis API, host-side synthesis, network TTS, external sound device, accelerated CPU, or $C030 speaker-based speech. Do not modify the emulator, ROM, bridge, assembler or shared checks. Report suspected platform defects separately. The browser may play normal emulator PCM, not process text or synthesize your voice.
Originality and permitted reuse
Create your own pronunciation rules, synthesis engine and voice data. General phonetics, DSP ideas and hardware documentation may inform the design; cite significant references.
Existing 3RIC hardware-access helpers may be reused or adapted: AY writes, initialization, timers, keyboard and text output. Identify reused routines and preserve applicable notices.
Do not copy, port or transliterate an existing TTS engine, pronunciation rules, voice tables or recordings. Do not use another TTS system or pretrained speech model to generate assets. No human voice recordings, canned words/sentences or precomputed demonstration/evaluation audio.
Original small synthetic waveform/phoneme tables are allowed. Include their generator and parameters if made offline, and include the generated assembly data in the source. A small original pronunciation-exception table may supplement general rules, not replace general text-to-speech with a fixed vocabulary.
Required application
Speak newly supplied ordinary English words and word sequences. Best-effort pronunciation is acceptable; routinely spelling words or silently omitting unfamiliar words is not a substitute for TTS.
Provide a text-mode UI with visible input, deletion and up to 120 characters. Enter speaks; completion returns to input for another sentence.
Required input is letters A-Z, spaces, apostrophes, hyphens and . , ? !. Normalize lowercase ASCII in the reusable API. Treat boundaries and punctuation sensibly. Numbers, abbreviations and non-English text are optional.
Handle empty/space-only input without speech or hanging. Report input limits instead of overflowing or silently losing part of the sentence. Reject or clearly explain unsupported characters.
Escape during speech cancels and silences sound, returning to input. Escape at input exits with BRK. Leave the keyboard/text output usable, no stuck sound, and no stray application-owned IRQ.
Separate UI from the reusable engine using the ABI below. A blocking engine is enough; concurrent game audio is not required.
Prioritize intelligibility and correct runtime behavior. Pitch/rate/voice controls are optional after the required app works.
Common callable interface
Export these case-insensitive assembler labels so shared checks and other programs can call the engine without an implementation-specific host adapter:
TTS_INIT: a subroutine callable immediately after loading the PRG, without first running the interactive UI. Establish required banks, tables and chip state. Return normally with RTS; do not wait for input.
TTS_SPEAK: a subroutine taking a pointer to NUL-terminated ASCII in RAM, low byte in A and high byte in X. Preserve the input string. Return with RTS, A=0 for completion (including empty input), A=1 for cancellation or A=2 for invalid input. Detect overlength input using the 121st character without unbounded scanning.
TTS_INPUT: reserve 122 consecutive bytes of always-mapped application RAM for callers/tests to supply text, including the overlength check and terminator. The application may use the same buffer. Do not overlap it with code or tables.
Keep the two callable entry labels in the loaded, always-mapped image. They may delegate to banked code. Return with upper ROM visible and sound silent. Document clobbered registers/flags, scratch, mappings, IRQ requirements and timing.
Shared checks allow 2 emulated seconds for initialization/empty/invalid calls, 90 seconds per spoken sentence, and at most 1 second to return after Escape. These are runtime safety bounds, not a development-time allowance or a quality score.
Memory, assembly and delivery
Use the project's asm6502.mjs dialect, not ca65 or invented directives. Symbols are case-insensitive.
Deliver one self-contained tts.s in your own submission directory, assembling both in Node and the existing browser workbench. Include entry.json and README.md.
Start at .org $0800 with startup code in always-mapped RAM. Use lower RAM through $BFFF by explicitly selecting RAM in the BASIC window when needed.
Banking is allowed; the entire 64 KiB space is not flat RAM. Document all regions, mapping transitions and any language-card ROM/interrupt safety strategy.
Establish bank state explicitly. Browser bulk loading fills backing memory directly; that is not proof physical BRUN loaded a bank. Include guest-side staging/relocation if needed and document a reproducible physical loading procedure.
Respect live ROM variables, vectors, stack and display. Document scratch and restore shared state when needed.
Emit a raw headerless tts.prg, load/entry $0800. The usual physical launch is BRUN TTS.PRG 0800; explain any additional bank-loading preparation.
No runtime assets downloaded or input-dependent preprocessing on the host.
State synthesis update/sample rate and memory/cycle budget, separating estimates from measured results. Keep addresses and clock assumptions identifiable for a port.
An Apple II port is not required. WOZ export does not prove Apple II compatibility.
Demonstration and verification
Use these public sentences and make the results reproducible:
HELLO. THIS COMPUTER CAN TALK.
WE MAKE NEW SOFTWARE FOR OLD COMPUTERS.
THE QUICK BROWN FOX JUMPS OVER THE LAZY DOG.
PLEASE TYPE A SENTENCE AND PRESS ENTER.
The host/viewers will also supply new ordinary English after the edition is frozen. Process it at runtime without changing source.
Run node codegen\tools\check-tts.mjs --entry <entry-id> after building the WASM runtime. Add your own tests for UI editing, input limits, repeated utterances, bank visibility, cancellation, silencing, IRQ cleanup and a usable monitor return. Include exact commands for generators and additional tests/recordings.
The shared checker assembles, exercises the callable ABI and captures normal emulator PCM. It does not establish intelligibility, AY-only compliance, interactive UI correctness, real ROM/SD loading or physical-hardware success. These need separate inspection/listening/testing. Nonzero PCM or plausible phoneme labels are not proof of speech. The generic harness.cjs treats WAI as a halt; do not use that as speech proof.
Capture the unmodified emulator running the actual 65C02 program. Host code may collect PCM, never generate or improve speech. Report sample rate and use normal playback speed. For browser listening, use native 1x and activate sound with a user gesture.
Commit and submit
Commit source, metadata, documentation, generators and tests in your entry directory. Generated PRGs/WAVs belong in ignored build output and downloadable check artifacts, not large binary commits. Record model identity/settings when known, agent environment, assistance, references and explicit licensing for new source/data.
Open an independent pull request to ebadger/3ric, using the repository's template. Do not overwrite another entry, modify shared checks/rules, push directly to main, self-merge or deploy. Follow required reviews; preserve the original candidate commit and identify any later human/second-model-assisted edition.
Return the PR URL, evaluated commit, reproduction instructions and an honest list of what was run, captured, heard, hardware-tested or left incomplete. If credentials prevent pushing/opening a PR, supply the local commit or patch and state the blocker. Do not claim a PR, review, test or published app that does not exist.
Human listeners make the main judgment: Can we understand it, and would we want this voice in our own 3RIC software? Leave something people can inspect, rebuild and improve.
Repository submission / maintainer review before publication
Submitting a 3RIC Talks entry
Read the complete v1 prompt before implementing. Challenge ID: tts-v1. Baseline commit: bd795bee12c5fb73a7de9e733e19e0aad95fb57c. The baseline is the commit that first added this version's prompt, not the latest main. Official entrants receive the same prompt, baseline, tools and host-specified limits. Community remixes are welcome; disclose their derivation and assistance.
Start from the shared baseline
Use a separate clone/worktree or a dedicated session starting at bd795bee12c5fb73a7de9e733e19e0aad95fb57c. Create a feature branch unless the session already supplies one. Never switch a shared working tree underneath another session. Do not read another entrant's implementation as a shortcut for an official same-prompt run.
Install/use Emscripten 6.0.1 as described in the web instructions. Node 22 or newer runs the tools. From the repository root, build the shared runtime:
pwsh -File web\build.ps1
Your entry directory
Choose a lowercase, hyphenated ID of at most 60 characters, such as my-voice-run-1. Create only your own directory under codegen\challenges\tts\v1\submissions\. Do not edit shared manifests, the gallery, platform code, this brief or the common checks.
codegen\challenges\tts\v1\submissions\my-voice-run-1\
tts.s
entry.json
README.md
test.mjs (optional additional tests)
generate.mjs (if you generate original tables)
The assembly must contain the common TTS_INIT, TTS_SPEAK and TTS_INPUT symbols specified in the prompt. Your README is the entry's implementation record: controls, callable ABI, memory/bank map, hardware loading, synthesis design, references, licenses, test commands and known shortcomings.
Start entry.json with this shape, replacing all example values with actual facts:
{
"schemaVersion": 1,
"challenge": "tts-v1",
"id": "my-voice-run-1",
"title": "My original voice",
"author": "your-name",
"model": "exact model ID/settings, or explicitly unknown",
"agent": "coding agent and tools used",
"baseline": "bd795bee12c5fb73a7de9e733e19e0aad95fb57c",
"description": "What this engine does.",
"assistance": "None, or a description of human/other-model help.",
"license": "License for new source and original data; preserve reused notices.",
"memory": "Actual regions, scratch and bank mappings.",
"limitations": "Known gaps and what has not been verified."
}
All these fields are required. Do not put credentials or private transcripts in metadata. Do not add self-awarded verification badges. Tests/recordings are evidence with a scope, not certification of originality, intelligibility or physical compatibility.
This validates metadata and assembly, calls your guest engine in the unchanged WASM VM with bounded execution, and writes PRG/WAV/report artifacts to codegen\out\tts-v1\my-voice-run-1\. Use --validate-only for a fast metadata/assembly check, not as a substitute for runtime checks. --all checks all entries.
The common checker does not run entrant-supplied JavaScript. Run and describe your own tests separately. It cannot certify the interactive application, bank loading through physical SD/ROM, AY-only sound, speech intelligibility or real hardware. List those results independently, and mark untested areas honestly.
To try the entry in the workbench, regenerate the challenge staging:
Open http://localhost:8011/index.html?src=programs/tts-v1-my-voice-run-1.s. Use 1x speed and a pointer/key gesture to enable audio. The normal editor supplies source inspection and PRG/WOZ export; WOZ is not evidence of Apple II compatibility.
Open an independent PR
Commit your source/metadata/README and any reproducible generators/tests. Do not commit toolchains or codegen\out artifacts. Follow the existing PR template and review rules; include challenge/entry ID, model/run information, evaluated commit, commands/results, audio evidence and physical-hardware status.
Check any existing PR's state before pushing. Never self-merge. If push permission is missing, provide a local commit/patch for the owner and say what remains blocked. The owner may need to approve Actions for a first-time contributor's fork.
The PR workflow uploads build/recording reports without production credentials. It does not publish an open PR to Pages. The owner reviews and merges before the entry appears on the public site. A failing entry stays inspectable in its PR.
Preserve the candidate commit used in the episode. If a reviewer, person or different model improves it, disclose that later edition rather than silently changing the claimed original result. Keep award/listening records attached to the evaluated revision.
Publication and judging
Merged entries get source, workbench and metadata links. The site does not infer speech quality or hardware success from a check result. The host will add listening/episode and People's Choice information when there are real entries and a ballot.
No submission fee, account requirement for playing, cloud synthesis service or automated model ranking is introduced by this challenge.