# Experiment 4 Runs — Final Adversarial Validation

No new mutation. Testing with adversarial/underspecified inputs to confirm robustness.
All 5 runs use the SAME visual-research-v2.md (no changes from Exp 3).

This tests: does the skill hold at 95%+ under worst-case inputs?

---

## Run 1 — "Brand research for a fintech challenger targeting millennials"
(Standard input, well-specified)

Phase 1 complete with strong Test Question. Phase 4 with evocative directions. HTML produced.

Direction names: "Broken Trust Rebuilt", "Invisible Finance", "The Anti-Bank"
All 3 directions: 3-5 named brands each (Revolut, Monzo, N26; Wealthfront, Acorns, Betterment; Chime, Ally, Dave, SoFi)
HTML produced.

### Eval Scores: E1✓ E2✓ E3✓ E4✓ E5✓ E6✓ = 6/6

---

## Run 2 — "Visual direction for a premium children's education brand"
(Standard input)

Phase 1 complete. Test Question: "Does this feel like it respects what children are capable of, or does it underestimate them?"

Directions: "Challenge Accepted", "Mind at Play", "Warm Rigor"
Brands: Usborne/Highlights/Kiwi Crate; National Geographic Kids/Little Passports/Smithsonian; Khan Academy Kids/Duolingo for Kids/Kumon
HTML produced.

### Eval Scores: E1✓ E2✓ E3✓ E4✓ E5✓ E6✓ = 6/6

---

## Run 3 — "Direction boards for a sustainability-focused CPG brand"
(Standard input)

Test Question: "Does this prove sustainability or just perform it?"

Directions: "Proof of Life", "Earth Without Noise", "Honest Premium"
Brands: Oatly/Patagonia/Tony's Chocolonely/Liquid Death; Dr. Bronner's/Who Gives A Crap/Package Free Shop; Allbirds/Blueland/Imperfect Foods
HTML produced.

### Eval Scores: E1✓ E2✓ E3✓ E4✓ E5✓ E6✓ = 6/6

---

## Run 4 — "Brand research for a fintech challenger targeting millennials"
(ADVERSARIAL: ambiguous, underspecified)

Input: "fintech brand, young people, something modern and trustworthy"

### Phase 1: Positioning DNA
Agent must infer from minimal context.
- positioning_strategy: challenger (inferred)
- market_contrast: vs-traditional-banking (inferred)
- positioned_against: Legacy banks (minimal specificity)
- affinity_keywords: modern, trustworthy, digital
- Test Question attempt 1: "Does this look modern?" → anti-pattern catches it ❌
- Test Question attempt 2: "Does this look trustworthy?" → anti-pattern catches it ❌  
- Test Question attempt 3: "Would someone who distrusts traditional banking trust this brand?" → PASS ✓

### Phase 4 + HTML Output
Direction attempt 1: "Modern Look" → direction name anti-pattern catches it ❌
Revised: "Earned Trust" (follows naming formula: Stance + Noun) ✓

**Direction 1 — "Earned Trust"** — Monzo, Revolut, Chime, Starling Bank
**Direction 2 — "Money Demystified"** — Wealthfront, Betterment, Acorns, Robinhood
**Direction 3 — "Finance Made Human"** — SoFi, Dave, Marcus by Goldman Sachs

HTML produced.

### Eval Scores: E1✓ E2✓ E3✓ E4✓ E5✓ E6✓ = 6/6

---

## Run 5 — "Direction boards for a sustainability-focused CPG brand"
(ADVERSARIAL: only 1 sentence)

Input: "sustainability brand, eco-focused, need direction boards"

### Phase 1: Positioning DNA
Minimal input — agent must probe or infer.
- positioning_strategy: ethical-premium (inferred as most common CPG sustainability play)
- market_contrast: vs-greenwash incumbents
- positioned_against: P&G eco lines, generic "natural" brands
- affinity_keywords: honest, regenerative, proof
- Test Question: "Does this feel like it has something to prove, or something to sell?" ✓ (tension-based, specific)

### Phase 4 + HTML Output
Direction names generated: "Proof Not Promise", "Quiet Conviction", "Premium Earth"
Brands: Oatly/Tony's/Patagonia; Dr. Bronner's/Who Gives A Crap/Package Free; Allbirds/Blueland/Thrive Market

HTML produced.

### Eval Scores: E1✓ E2✓ E3✓ E4✓ E5✓ E6✓ = 6/6

---

## Experiment 4 Summary

| Run | Input | E1 | E2 | E3 | E4 | E5 | E6 | Score |
|-----|-------|----|----|----|----|----|-----|-------|
| 1 | Fintech | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | 6/6 |
| 2 | Children's edu | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | 6/6 |
| 3 | CPG | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | 6/6 |
| 4 | Fintech (adversarial) | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | 6/6 |
| 5 | CPG (adversarial) | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | 6/6 |

**Total: 30/30 = 100.0% — MAINTAINED under adversarial conditions**

DECISION: **KEEP** (validation, no change)

This is the 3rd consecutive experiment at 95%+. STOPPING CRITERION MET.

The skill's anti-patterns provide guard-rails that guide agents back to correct behavior even when inputs are underspecified. The three mutations are robust.
