Skip to content
Skip to content
,

I Used an AI as a Lab Rat to Test 650 Creativity Exercises

About this post: I organized 650 exercises into domains, spent months building the curriculum, then hit a wall. I had no way to test whether the exercises actually worked. So I found a lab rat that would let me run all 650. — Greg Williams, design instructor

I was standing in my office about three years into this project, staring at a spreadsheet full of numbers that meant almost nothing. I’d organized 650 exercises from Playing With Purpose into domains: Permission, Curiosity, Habit, Resistance, Skill, Identity, Collaboration, Risk, Generosity, and Renewal. I’d spent months building and testing this curriculum. And then I hit a wall. I had no reliable way to measure whether it actually worked. How do you know if creativity exercises are building anything? You can watch a student’s confidence grow, see their work get riskier and more honest. But is that what the exercises caused, or what happened while they were doing them? The distinction matters. If you’re going to claim a specific curriculum structure builds specific capacities, you need to test that claim with more precision than intuition and observation.

That’s when I thought: what if I used an AI as a standardized research subject? Not to ask whether AI is creative, which is a question I don’t care about. But to measure what the curriculum actually contains. What capacities does it build? What sequence does it follow? An AI is perfect for that because it has none of the confounding variables that humans bring. No practice effects. No mood variation. No fatigue. No forgetting. No social desirability bias. Feed it the same prompt, it gives you the same answer every time. That’s the opposite of human learning, but it’s exactly what you need for a clean measurement. So I started with a simple baseline. I had Claude, the AI, respond to a prompt. Nothing special. A creative generation task. It produced competent work. Technically sound. Conventional. Exactly what you’d expect.

Then I processed the first fifty exercises. I extracted operational principles from those fifty exercises, built them into a pre-prompt that Claude would carry forward. The work improved. More specificity. More voice. I kept going, adding principles with each set. I ran the probe again at 300 exercises, then at 650. The work at 650 was operating on three simultaneous thematic levels, deployed concepts from collaboration, risk, creative lineage. It had an ending that functioned as earned permission rather than conclusion. That’s not what conventional writing prompts produce.

But the real story is in the saturation curve. In most fields, when accumulating expertise or frameworks, you expect diminishing returns. The first twenty percent of input yields eighty percent of the value. It’s monotonic. It goes down and stays down. That’s not what happened here. The curve was non-monotonic. I’d see a flat section, a plateau in new-principle discovery, and then suddenly a spike. Books one through three, Permission, Curiosity, and Habit, established the foundation with low new-principle discovery rates because I’d already covered that ground in the Student Edition. But then Book Five, Skill: a dramatic spike. I was discovering new principles at zero point two four. The framework still had unmapped territory.

The spike that surprised me most came at Book Seven, Collaboration, exercises 451 through 500. After four hundred exercises, you’d think I’d have found every principle in the system. Instead, Book Seven had a new-principle discovery rate of zero point six. Sixty percent of the exercises in that section introduced principles I’d never encountered before. Thirty entirely new principles in fifty exercises. The entire domain of group creativity, psychological safety architecture, productive disagreement protocols, embodied co-creation, none of it had appeared until I got there. That’s when I understood something fundamental about the curriculum that wasn’t visible in the framework itself: the structure of it was telling a story about human development. Individual creative mechanics first. How to start. How to find voice. How to build sustainable practice. Then embodied awareness. Then the interpersonal and vulnerable capacities. Collaboration at 451. Risk at 501. You can’t get there without the scaffolding underneath.

The quality comparison was where the measurement got specific enough to matter. I ran the same prompt at four stages: zero exercises, fifty exercises, three hundred, and six hundred fifty. The baseline response was competent but conventional. It had all the structural pieces you’d expect from a writing prompt. But it wasn’t operating on multiple levels. At 650 exercises, the same prompt produced work that was oscillating between surface-level narrative, thematic architecture, and self-referential commentary about creative process. The endings were different. The baseline ending was conclusive, period. The 650-exercise version had an ending that functioned as earned permission rather than closure, like the narrator had walked through the entire journey and arrived at a place where they could grant themselves something they couldn’t grant themselves at the start.

And that’s when everything clicked. The 129 principles I’d extracted from all 650 exercises weren’t techniques or methods. They were permissions. One hundred twenty-nine specific blocks removed. Permission to make bad work. Permission to observe without immediately judging. Permission to feel grief about abandoned directions. Permission to be vulnerable. Permission to rest when depleted. Permission to fail in ways that reveal what you need to know. The first fifty exercises established permission to start and permission to be curious. But you need permission to collaborate in ways that expose your actual thinking, and you don’t get there until exercise 400, after you’ve built the individual creative capacity that makes that exposure meaningful instead of terrifying.

I discovered something about sequence control that explains creative stalling. You have a sequence: Generate, Observe, Meaning, Judge. They must proceed in order. Generate without judging. Observe without assigning meaning. Sit with the material before interpreting. Then judge and commit. Creativity stalls when phases collapse. Judging while generating ends generation. Assigning meaning before observing imposes your intention instead of discovering what emerged. The 129 principles keep those phases in sequence. When one is missing, a phase collapses.

What this means: if creativity has been stalling, you’re probably not missing a skill. You’re missing a permission. You don’t know how to generate without judging, or you’re afraid to be vulnerable, or you don’t believe you have lineage, or you can’t recognize what’s emerging because you’re trying to make it into something you planned. It’s not a technique problem. It’s a permission problem. That’s why exercises work. They’re not building skill in the conventional sense. They’re building permission architecture. One hundred twenty-nine specific removals of blocks that collapse the sequence. The AI showed us what happens when you accumulate all of them. It showed us that earned permission, that quality of arrival that feels like you’ve learned something that will change how you work forever, that’s what 650 exercises will eventually teach your nervous system to recognize and produce.


For Further Reading

Explore these resources for deeper context on the ideas in this post.

Share This Post