A hands-on experiment exploring creativity, judgement and what changes when AI becomes part of the design process
The question wasn't whether AI could build a product. It was whether I could still recognise good design when it did.
During 2025, I've watched AI move from curiosity to everyday practice. Designers are using it to generate concepts, write copy, explore ideas and accelerate production. Whether organisations have formal guidance or not, AI has already become part of many design teams.
As a design leader, I found myself asking a different question. How do you lead a team through this well? Not by reading think pieces or watching demos, but by experiencing the messiness firsthand.
So I set myself a challenge. I gave myself 48 hours, a blank brief and one rule: Build a product with AI as a genuine creative partner, while remaining responsible for every important decision. The result was Tiletopia, a small daily sliding puzzle game.
The game wasn't the goal.
Understanding the changing role of design leadership was.

I deliberately chose a product that was simple enough to build quickly but rich enough to expose real design decisions.
From the beginning, I wanted Tiletopia to feel calm, nostalgic and quietly rewarding. The kind of experience you'd return to for five minutes each day because it left you feeling a little better than before. I wasn't interested in shipping a polished product.
I wanted a realistic creative process where ideas evolved, assumptions were challenged and AI became something more than a content generator.
It quickly became clear that the process was far messier than I'd imagined. Looking back through my conversations with ChatGPT, the journey wasn't linear at all. More than thirty conversations. Concepts abandoned. Features removed. Scope reduced. Moments where I typed "change of plan" because I realised I was building too much before proving the experience itself.

It accelerated the parts of design that benefit from speed and exploration.
AI helped me:
• Explore different gameplay mechanics.
• Generate alternative interaction patterns.
• Pressure-test ideas quickly.
• Iterate on flows and reward systems.
• Surface possibilities I hadn't initially considered.
What would previously have taken hours or days could often be explored in an afternoon. The speed was genuinely powerful. But it also revealed a critical leadership challenge.
AI accelerates decisions.
While AI accelerates good decisions. It also accelerates bad ones.
It doesn't improve judgement.
The biggest lesson from the experiment wasn't about AI. It was about leadership. AI will happily generate ten reasonable solutions before you've fully understood the problem to identify the best decision. While that speed is exciting. It's also dangerous. Without clear principles, AI doesn't simply accelerate creativity. It accelerates drift.
Several ideas looked perfectly sensible until they collided with the experience I was actually trying to create. Which made me realise something. As AI becomes more capable, the value of design leaders shifts even further away from producing outputs. It shifts towards protecting intent.
Several early directions looked reasonable but were wrong for the experience I wanted to create.
For example, AI suggested:
• More complex grid systems to increase difficulty
• Multiple images themes to increase variety
• Emoji-heavy interfaces to make the experience feel loud, chatty and frantic.
Individually, these suggestions made sense. Together, they moved the product away from its intended feeling.
The experience was supposed to be calm, nostalgic and discovery-led. The AI-generated direction was becoming louder, more complex and less human.
The issue wasn't that AI was wrong. The issue was that AI could not understand the emotional intent behind the product unless I provided the judgement and context.
The most interesting moments were not where AI succeeded. They were where it reached its limits, and also where it couldn't replace lived experience.
First example:
AI recommended sports as the strongest theme for the first version because images were easy to recognise and translate into puzzles. Logically, it was a good recommendation. But it missed something important. The intended end experience for a user was not about solving a puzzle. It was all about discovery.
I understood that world landmarks created that sense of curiosity and connection: revealing a place tile by tile, discovering something new, and creating a reason to return. That decision came from my personal judgement.
Another example:
Multiple times I asked AI for a modern, timeless visual direction. What came back wasn't bad. It just wasn't right. Some concepts felt dated. Others felt overly clinical. None captured the feeling I wanted.
Then I remembered a trip to Portugal. Walking through Lisbon, I'd become fascinated by Azulejo tiles; geometric, handcrafted and quietly beautiful. For a game called Tiletopia, the connection suddenly felt obvious.That became the visual foundation for the product. Not because AI discovered it. Because I brought a memory into the process.
That single decision shaped everything else. The language shifted from "Today's Puzzle" to "Today's Discovery."
Success became "Discovery Unlocked" instead of simply "Completed." Small changes, but they completely changed the emotional tone of the experience. The game became about uncovering places.
Last example:
Sound was the most persistent struggle. The first background audio AI suggested felt like an alarm. When I asked for relaxation it swung so far the other way it felt like a meditation app for people already asleep. Then when I asked AI to choose a royalty-free calm track, it listed sources and told me to pick one myself.
AI could describe what I needed, but it still couldn't make the judgment call about what actually felt right.
Rather than forcing every player into the same interpretation of "calm", I added a simple audio toggle instead. Ironically, failing to solve the problem produced the better design decision.
In summury AI could remix patterns, but it could not bring the experiences, emotional connections or moments that gave those patterns meaning.

There was one moment where AI genuinely expanded my thinking. Without prompting, it suggested numbering puzzle tiles to support players with visual impairments. I hadn't considered it.
The suggestion made me rethink an area I'd unintentionally deprioritised. But it also reinforced something equally important. The feature wasn't tested. It wasn't validated, so I couldn't confidently claim it improved accessibility simply because AI suggested it.
AI can broaden thinking. It cannot replace responsibility. This is the tension I found the hardest to resolve: while AI can open up thinking you wouldn't have done yourself, it can also create a false sense that something important has been handled.
The most uncomfortable part of the experiment wasn't where AI failed. Instead it was where I stopped questioning it.
There were moments where AI generated screens that looked polished enough for me to accept without much thought. Not because they were excellent. But because they looked familiar.
Only later, when I deliberately changed things like the horizontal progression, blurred undiscovered landmarks and simplified the visual hierarchy, did I realise what had happened. I'd quietly moved from designing to curating. That distinction has stayed with me ever since.
The danger isn't that AI produces bad work. It's that it often produces work that's good enough to stop us asking whether it's actually the right answer. As a design leader, I think that's one of the biggest risks we'll face.
The biggest outcome wasn't the game. It was a shift in how I think about design leadership. Design leaders become guardians of judgement. As AI accelerates production, the real differentiator becomes knowing what should be built, what should be challenged and what should be protected.
Teams need AI fluency, not AI dependency
The question isn't whether teams should use AI. They already are.
The challenge is helping them understand where AI creates leverage, where it introduces risk and where human judgement remains essential.
Principles matter more than prompts
Without clear product principles, AI naturally fills the gaps with statistically likely answers. Fast becomes familiar. Familiar becomes average. Leadership increasingly means creating the conditions where originality can still emerge.
Trust becomes a design responsibility
As AI-generated content becomes more convincing, people will question what is authentic. Designers will need to play a critical role in protecting trust, transparency and confidence and also creating relevant experiences.
That's not just a design problem. It's a leadership one.
When I started this experiment, I thought I was exploring AI. Looking back, I was really exploring myself.
The game still isn't finished. Retention wasn't as strong as I'd hoped. The soundtrack still isn't quite right. Some ideas didn't survive contact with real users. And that's exactly why the experiment was valuable. It reminded me that AI can generate options. It can accelerate iteration. It can challenge assumptions. But it can't replace the experiences, judgement and curiosity that shape meaningful products.
The future of design isn't humans versus AI. It's designers who know how to work with AI while protecting the things that make great design human. For me, that's where design leadership becomes even more important.
Beta is still ongoing: 5 of 20 users over 7 days. Happy to add more testers.