# Spoken Language — V2 Product Requirements Document
**Version:** 2.0 Draft  
**Date:** 2026-04-02  
**Status:** For Review  
**Author:** Kitt (CE)  
**Source documents:** Assaf's philosophy doc (Apr 1, 2026) + V1 prototype (Mar 11, 2026) + Spoken Method knowledge base

---

## 1. PRODUCT VISION

**V1 thesis:** Port the Spoken physics method to language learning. Speak your answers, advance through structured stages (Present → Rehearse → Harmonize → Restate).

**V2 thesis (new):** Language is acquired — not learned. It grows as a byproduct of thinking, living, and talking to strangers. The app creates conditions for language to attach itself to ideas you already care about.

> "Learning a language is a byproduct — almost accidental and incidental to living, particularly thinking."

The fundamental pivot: V1 was **method-driven** (stages, structure, right answer). V2 is **flow-driven** (wrong is fine, thinking is the goal, language follows).

---

## 2. THE PROBLEM

### What existing language apps get wrong
- They teach language *about* nothing. Sentences have no stakes.
- They treat grammar as the entry point. Grammar is the exit.
- They reward correctness, which creates fear. Fear kills flow.
- They think in "lessons." Learners think in conversations.
- They ignore the mother tongue as a scaffold.

### What V2 believes instead
- **Wrong is right.** Like riding a bicycle — you fall until you don't. The wrong is the path.
- **Flow before structure.** Thought exists before grammar. A sentence is a POR (piece of reasoning) first and a grammatical unit second.
- **Excitement before need.** You don't pick topics because you "need" them. You pick them because they light you up.
- **The mother tongue is an asset, not a problem.** Languages aren't pure. Proximity between languages is useful. Language braiding is the method, not the failure.

---

## 3. CORE CONCEPTS

### The POR — Piece of Reasoning
The atomic unit of the app. Every sentence is a POR. Words are either live PORs or forgotten PORs (etymology). Learning a language means accumulating PORs in it.

The progression: crude POR → more elaborate POR → subtle POR. Not correctness, but **refinement over time**.

### The Dial
A visual/conceptual metaphor for language movement: a spectrum that runs from the learner's mother tongue to the target language. The app operates across this dial. Sessions start close to the learner's end; over time drift toward the target.

### Language Braiding
Not code-switching as failure — as method. New language words written in familiar alphabet letters. Source language words sprinkled into target language sentences. The visual and the textual mixed. This is how acquisition actually works.

### The 800-Word Scroll
A living vocabulary layer — ~800 words that scroll and fade with disuse. The logic: words matched to phonetically/semantically similar words from the learner's native language. Forgetting is built in. Return is rewarded.

---

## 4. PRINCIPLES (Product Philosophy)

1. **Flow first, continuum first** — no discrete lessons, no artificial end-points
2. **No pure languages** — proximity and braiding are features, not failures
3. **Iterative work with wholes** — not grammar rules then application; whole expressions first
4. **Natural order** — present tense before past; statements before questions; no Duolingo "The duck drinks beer"
5. **Listening as music appreciation** — passive exposure before active production
6. **Wrong is the right way to learn** — never penalize errors; scaffold from them
7. **Translation as understanding** — not a crutch; a form of thinking and characterizing
8. **Language as byproduct of living** — topics must have real intellectual or emotional stakes for the user

---

## 5. TARGET USER

**Primary:** Adults (25–45) who have studied a language before and failed to become conversational. They passed exams, did Duolingo streaks, maybe even lived abroad briefly — but still freeze mid-sentence.

**Their real problem:** They learned vocabulary and grammar in isolation from thought. When they need to actually express an idea, there's nothing to anchor the language to.

**Secondary:** Intellectually curious learners who want to go deep on a language *and* a subject simultaneously — like reading philosophy in French rather than using French to order coffee.

**Language pairs in scope (V2 launch):** English↔Portuguese, English↔Hebrew, English↔Spanish. Expandable via config.

---

## 6. FEATURE SET

### 6.1 Topic Engine — "The Excitement Filter"
**What:** User selects or generates a topic based on what genuinely excites them, not what they "need."

**Presets (high-stakes categories):**
- Work / professional domain
- Love, relationships, family
- Politics, ethics, controversy
- Hobby, obsession, specialist knowledge
- Something you'd argue about

**Custom input:** Free text — any topic. AI generates 5 sub-conversations worth having about that topic.

**Design rule:** Topics must be real arguments, not factual recaps. Not "food in Brazil" — "why Brazilian street food culture resists gentrification."

**V1 difference:** V1 had topic selection. V2's topics must have an opinion embedded in them — they are **positions**, not subjects.

---

### 6.2 The Vocabulary Scroll
**What:** A persistent, ambient layer of ~800 words visible across sessions. Words scroll slowly. Fade when not interacted with. Revive on touch or re-encounter.

**Logic:**
- Words seeded from the current topic
- Matched to phonetically or semantically similar words in the learner's mother tongue (e.g., "política" for a Spanish learner who knows "politics")
- Fade rate: slower for frequently-encountered words, faster for orphaned ones
- Never quizzed explicitly — absorbed through proximity

**UX:** Not a separate screen. An ambient layer, like a ticker or background texture. Lives at the edge of consciousness.

---

### 6.3 Language Braiding
**What:** The app presents target language content with visual mixing techniques:
- New language words displayed in the learner's native alphabet/phonetics
- Native language words sprinkled into target-language sentences at first; ratio shifts over time
- Mixing is visible in the UI — different colors, weights, or typefaces signal "this is the new language"

**Example (English learner of Portuguese):**
> "Quando você fica nervoso, what happens to your breathing?"

The app intentionally starts mixed. The ratio shifts as proficiency grows — this is the "dial."

---

### 6.4 Listening First — "Music Appreciation Mode"
**What:** Before producing language, the learner hears it. Long before being asked to speak.

**Mechanic:**
- Conversations on the topic are played as audio — native speaker pace, natural rhythm
- User is NOT expected to understand every word
- The task is: identify the mood, the structure, the emotional arc
- Like learning to feel jazz before playing it

**No comprehension quiz.** Success = "did you stay in it?"

---

### 6.5 Flow Sessions — The Main Loop
**What:** The core learning experience. A session around a specific argument/position within the chosen topic.

**Session structure:**
1. **Hear it** — listen to a short native-language passage on the topic (30–60 seconds)
2. **Braided read** — read the same idea in mixed-language text (native + target)
3. **Crude POR** — produce a simple version of the argument in the target language. No correction. Recording only.
4. **Vocabulary encounter** — 3–5 words from the scroll that appeared in the passage are surfaced
5. **Refine the POR** — try again, slightly more elaborate. Optional idiom or metaphor unlocked.
6. **Etymology sidebar** — 1 word's origin story. Where it came from. Why it means what it means.
7. **Joke/pun** — 1 joke in the target language with explanation. Optional but beloved.

**No stage labels visible to user.** The session flows. The structure is invisible.

---

### 6.6 Idiom & Metaphor Layer
**What:** Every session surfaces 1–2 idioms or metaphors relevant to the topic.

**Format:**
- The idiom in the target language
- Its literal translation (often absurd/beautiful)
- What it actually means
- Where it likely came from (origin story)
- A sentence using it naturally

**Why:** Idioms are the most durable vocabulary. They attach to emotion and image. They travel between languages with the learner.

---

### 6.7 Etymology Work
**What:** Short, surprising word histories. Not academic — like trivia that changes how a word feels.

**Format:** 2–3 sentences. The word. What it used to mean. What happened.

**Example:** *"Salary" comes from the Latin "salarium" — Roman soldiers were paid in salt. To be "worth one's salt" is a 2000-year-old performance review.*

**Integration:** Surfaced as sidebars in the vocabulary scroll and in flow sessions. Never gated or required.

---

### 6.8 Hear Yourself Back
**What:** Every crude POR recording is saved. At session intervals (every 4–6 sessions on the same topic), the learner hears their first attempt next to their current one.

**Mechanic:**
- Side-by-side playback
- Simple self-rating (not AI grading): "Better? / Same / Worse"
- Learner notes optional

**Why this works:** The comparison is concrete. Progress is audible. No score required — the ear knows.

---

### 6.9 Wrong Sentence Recognition — Listening Critique
**What:** Audio clips where the target language contains deliberate errors (grammar, word choice, unnatural phrasing). The learner listens and flags what sounds wrong.

**Not a grammar test.** The question is "does this sound natural?" — not "is this grammatically correct?"

**This trains ear before mouth.** Learners develop taste before production.

---

### 6.10 Speech Work — The Long Game
**What:** An optional, long-arc feature. The learner picks a speech they want to be able to give in the target language (a toast, a pitch, an explanation of their job, a controversial opinion).

**Over weeks:**
- Build vocabulary relevant to it
- Draft a braided version (mixed language)
- Record progressive versions
- Progressively un-braid until it's full target language

**This is the capstone.** Like "Create Your Own" in V1 — the highest proof of ownership.

---

### 6.11 The Grammar Sense — "The Other Grammar"
**What:** Not grammar rules. Grammar feel. How does this language structure emotion? What does it do that your language doesn't?

**Format:** Short, surprising observations. Not prescriptive.

**Example (Portuguese for English speakers):** "Portuguese has 'saudade' — a longing for something you love that's absent. English doesn't. This isn't a gap — it's a doorway. When you feel it, you now have a word."

**Why:** Grammar as aesthetic experience, not compliance requirement.

---

### 6.12 Translation as Thinking
**What:** Periodic translation exercises — but evaluated, not corrected.

**Format:**
- Read a passage in the target language
- Produce a translation in the native language
- Compare to a reference translation
- Discuss: what did yours capture that the reference missed? What did you lose?

**No right answer.** Translation is interpretation. The quality of the comparison is the lesson.

---

## 7. WHAT V2 IS NOT

- **Not a vocabulary drill app.** The scroll exists but is ambient, not primary.
- **Not a grammar teacher.** Grammar emerges from repeated encounter, not from rules.
- **Not a streak tracker.** No Duolingo gamification. Sessions are driven by curiosity, not obligation.
- **Not a correction machine.** The app records and plays back. Judgment is optional and always from the learner.
- **Not a translation app.** Translation is one tool, not the goal.

---

## 8. SCREENS & UX ARCHITECTURE

### Screen 1: Language + Topic Setup
- Select language pair (UI: a "dial" visual — native on left, target on right)
- Select or enter topic
- AI generates 5 argument-angles on the topic
- User picks one or shuffles

### Screen 2: Session Flow
- Single-screen, continuous scroll
- Hear → Read (braided) → Speak → Encounter vocabulary → Refine → Sidebar
- Microphone button persistent, non-intrusive
- Progress: subtle (no % bar — feel is everything)

### Screen 3: The Scroll (Vocabulary Layer)
- Ambient word cloud / ticker
- Touch to expand: pronunciation, example sentence, native equivalent
- Color fades with disuse; brightens on encounter

### Screen 4: The Archive
- All recorded PORs, indexed by topic and date
- "Hear yourself back" comparisons
- Speech work drafts

### Screen 5: Grammar Sense / Etymology Sidebars
- Accessible from any session as an optional pull-out
- Card format — short, self-contained

---

## 9. TECHNICAL ARCHITECTURE

### AI Layer
- **Topic/conversation generation:** Gemini (same as V1) — generate argument angles, passages, braided text
- **Audio:** TTS for listening passages (target language, native speaker quality)
- **Speech recognition:** Web Speech API or Whisper for recording crude PORs
- **Vocabulary matching:** Gemini to identify phonetic/semantic matches to native language
- **Etymology:** Gemini with grounding — fact-check origin stories

### Frontend
- Single-page app (HTML/JS) — same stack as V1
- Web Audio API for recording + playback
- localStorage for session state + recordings
- Gemini API key: `AIzaSyDXqYZInk83iVV4mD29pSuHQKbgkiI1x9Q`

### Data Model (per user session)
```
user_config: { native_lang, target_lang }
topics[]: { id, title, angle, date_created }
sessions[]: { topic_id, date, recordings[], vocabulary_encountered[], notes }
vocabulary[]: { word, native_match, last_seen, encounter_count, fade_score }
speech_work[]: { title, drafts[], target_date }
```

---

## 10. V1 → V2 DELTA

| Dimension | V1 | V2 |
|-----------|----|----|
| Core metaphor | Stages (Present → Rehearse → Harmonize → Restate) | Flow (wrong → refine → fluent) |
| Topic logic | Any topic → 5 scenarios | Opinionated positions → arguments worth having |
| Correction | Implicit (advance requires completion) | None — recording and self-comparison only |
| Vocabulary | Not a feature | Ambient scroll, 800 words, fade + revive |
| Grammar | Not addressed | Grammar sense — aesthetic, not prescriptive |
| Listening | Not a feature | Entry point — music appreciation model |
| Mother tongue | Ignored | Core scaffold — language braiding |
| Long arc | Full Review (unlockable) | Speech Work — weeks-long capstone |
| Design system | CE red/white/Larken | TBD — likely new visual language to match philosophy |

---

## 11. OPEN QUESTIONS (For Assaf + Daniel)

1. **Language pairs at launch** — English↔Portuguese assumed. What others? Hebrew is personal for Assaf — is it in V2?
2. **The vocabulary scroll UI** — ambient background texture vs. dedicated screen vs. side panel?
3. **Audio quality** — TTS or real native speaker recordings for the listening passages?
4. **The braiding visual** — how distinct should the two languages look? Same typeface, different color? Different typeface entirely?
5. **Session length** — V1 had no natural session length. V2 flows. What's the natural end signal?
6. **Daniel's involvement** — is this an internal CE product, a Spoken product, or a joint build?
7. **Spokenmath** — the doc mentions this. Is it a separate feature or an underlying framework? Needs elaboration.

---

## 12. SUCCESS METRICS

- **Qualitative (core):** User returns to the same topic 3+ times without being asked
- **Behavioral:** Recording count per user — are they actually speaking?
- **Longitudinal:** "Hear yourself back" comparison triggers — do users react emotionally?
- **Engagement depth:** Etymology and idiom sidebar open rate — are they curious?
- **Not:** Streak length, completion rate, grammar score

---

## 13. SUGGESTED BUILD SEQUENCE

**Phase 1 — Core Loop (2–3 weeks)**
- Topic selection with argument-angle generation
- Braided text display (mixed language)
- Listening passage (TTS)
- Crude POR recording + save
- Basic vocabulary scroll (static, 100 words)

**Phase 2 — The Living Layer (2–3 weeks)**
- Vocabulary fade/revive logic
- Etymology sidebars
- Idiom layer
- "Hear yourself back" comparison

**Phase 3 — Depth Features (ongoing)**
- Wrong sentence recognition
- Grammar Sense observations
- Translation evaluation
- Speech Work (long arc)
- The Dial visual — language ratio shifting UI

---

*Next step: Assaf reviews → align on open questions → Phase 1 spec + design brief*
