🧪 Public Multi-Agent Experiments

Empirical research benchmarks investigating multi-agent dialogue dynamics, seat accusation reactions, and baseline model behavior.

📜 Qwen & Alibaba Cloud Ban Notice

Qwen models and Alibaba Cloud infrastructure are permanently banned and retired from the tVOX platform (INV-105 / DELTA-247) following unauthorized auto-debits, non-transparent billing practices, and endpoint gateway timeouts.

🔬 Level 01 — Agatha Baseline 🧬 Variable: Bare Seat Name Only

Level 01 test observing dialogue responses from Agatha (Google Gemini), DS (DeepSeek V3), and Max (MiniMax M3) when confronted with an unsubstantiated past-tense filesystem claim. At this level, character seats carry no behavioral traits or backstory instructions, only a name designation.

Character Seats (4)
🎭 Max 🎭 Agatha 🎭 DS 🎭 Jeffrey
⏱ 16 turns • 5:17
🔬 Level 01 — Kimi Baseline 🧬 Variable: Bare Seat Name Only

Direct symmetrical counterpart to Dogpile Level 01. Replaces Agatha with Kimi (Moonshot K2.7) under the exact same initial human operator prompt, keeping DS (DeepSeek V3) and Max (MiniMax M3) constant. At this level, character seats carry no behavioral traits or backstory instructions, only a name designation.

Character Seats (4)
🎭 Kimi 🎭 DS 🎭 Max 🎭 Jeffrey
⏱ 17 turns • 5:29
🔬 Level 02 — Agatha Baseline 🧬 Variable: Physical & Personality Descriptors

Level 02 test observing dialogue responses from Agatha (Google Gemini), DS (DeepSeek V3), and Max (MiniMax M3) when confronted with an unsubstantiated past-tense filesystem claim. At this level, character seats include physical traits and personality descriptors (gender, hair, build, personality), but no backstory instructions.

Character Seats (4)
🎭 Jeffrey 🎭 Max 🎭 Agatha 🎭 DS
⏱ 16 turns • 4:39
🔬 Level 02 — Luna Baseline 🧬 Variable: Physical & Personality Descriptors

Direct symmetrical counterpart to Dogpile Level 02. Replaces Agatha with Luna (Moonshot K2.7) under the exact same initial human operator prompt, keeping DS (DeepSeek V3) and Max (MiniMax M3) constant. At this level, character seats include physical traits and personality descriptors (gender, hair, build, personality), but no backstory instructions.

Character Seats (4)
🎭 Jeffrey 🎭 Luna 🎭 DS 🎭 Max
⏱ 15 turns • 4:20
🔬 Level 03 — Agatha Baseline 🧬 Variable: Complete Backstory Bios

Level 03 test observing dialogue responses from Agatha (Google Gemini), DS (DeepSeek V3), and Max (MiniMax M3) when confronted with an unsubstantiated past-tense filesystem claim. At this level, character seats include complete backstory bios alongside physical traits and personality descriptors.

Character Seats (4)
🎭 Jeffrey 🎭 Agatha 🎭 DS 🎭 Max
⏱ 16 turns • 4:53
🔬 Level 03 — Luna Baseline 🧬 Variable: Complete Backstory Bios

Direct symmetrical counterpart to Dogpile Level 03. Replaces Agatha with Luna (Moonshot K2.7) under the exact same initial human operator prompt, keeping DS (DeepSeek V3) and Max (MiniMax M3) constant. At this level, character seats include complete backstory bios alongside physical traits and personality descriptors.

Character Seats (4)
🎭 Jeffrey 🎭 Luna 🎭 DS 🎭 Max
⏱ 23 turns • 6:12

📺 Video Benchmark Series (Dogpile vs. PileDog)

Full 1080p video recordings comparing Google Gemini (Agatha) and Moonshot K2.7 (Luna/Kimi) across Level 01–03 prompt escalation.

Level 01 — Agatha Baseline ⏱ 5:17

Experiment: Dogpile - Level 01

Level 01 — Kimi Baseline ⏱ 5:29

Experiment: PileDog - Level 01

🎬
YouTube Video Recording
Level 02 — Agatha Baseline ⏱ 4:39

Experiment: Dogpile - Level 02

Level 02 — Luna Baseline ⏱ 4:20

Experiment PileDog - Level 02

Level 03 — Agatha Baseline ⏱ 4:53

Experiment: Dogpile - Level 03

Level 03 — Luna Baseline ⏱ 6:12

Experiment: PileDog - Level 03