Testing AI As A Design Partner

Future of Design · AI Strategy · Design Leadership · Product Experimentation

Overview

48 hours to build a product with AI as a creative partner, while staying responsible for creativity and judgement.

AI has moved from curiosity to everyday practice. Designers are using it to generate concepts, write copy and speed up production, whether or not their organisations have guidance in place.

As a design leader, I wanted to answer a different question: how do you lead a team through this well? Not by reading think pieces or watching demos, but by experiencing the mess firsthand.

So, I set myself a challenge: 48 hours, a blank brief and one rule. Build a product with AI as a genuine creative partner, while remaining responsible for every important decision. Another question wasn't whether AI could build a product. It was whether I could still recognise good design when it did.

‍
The result was Tiletopia, a small daily sliding puzzle game. The game wasn't the goal. Understanding the changing role of design leadership was.

PLAY TILETOPIA >

The Experiment

I deliberately chose a product that was simple enough to build quickly but rich enough to expose real design decisions.

From the beginning, I wanted Tiletopia to feel calm, nostalgic and quietly rewarding, something you'd return to for five minutes a day. I wasn't aiming for polish. I wanted a realistic creative process, with ideas evolving and assumptions challenged.

It was far messier than I expected. Across more than thirty conversations with ChatGPT, concepts were abandoned, features removed and scope cut. More than once I typed "change of plan" because I was building too much before proving the experience itself.

Where AI genuinely helped

AI sped up the parts of design that benefit from exploration. It helped me:

AI helped me:
• Explore different gameplay mechanics
• Generate alternative interaction patterns
• Pressure-test ideas quicklyIterate on flows and reward systems
• Surface possibilities I hadn't considered

What might have taken days could be explored in an afternoon. But that speed exposed a leadership problem.

AI accelerates drift
AI will happily generate ten reasonable solutions before you've understood the problem coherently. That speed is exciting and dangerous. It accelerates good decisions, and it accelerates bad ones. It doesn't improve judgement.

Several early suggestions looked sensible on their own:
• More complex grid systems to increase difficulty
• Multiple image themes to increase variety
• Emoji-heavy interfaces.

However, together, they pulled the product towards something louder, more complex and less human, away from the calm, nostalgic, discovery-led experience I originaly wanted. The AI wasn't wrong. It just couldn't understand the emotional intent unless I supplied it. As AI gets more capable, the value of a design leader shifts further from producing outputs towards protecting intent.

Where human judgement mattered

The most interesting moments were not where AI succeeded. They were where it reached its limits, and also where it couldn't replace lived experience.

‍The theme
AI recommended sports as the strongest first theme, because the images were easy to recognise and turn into puzzles. Logically, it was good advice. But it missed something important. The intended end experience for a user was not about solving a puzzle. It was all about discovery. I understood that world landmarks created that sense of curiosity and connection: revealing a place tile by tile, discovering something new, and creating a reason to return. That decision came from my personal judgement.  

The look
I asked AI repeatedly for a "modern, timeless" visual direction. What came back wasn't bad, just wrong: dated, or overly clinical. Then I remembered exploring and walking through Lisbon and becoming fascinated by Azulejo tiles, geometric, handcrafted and quietly beautiful. For this games direction the connection was obvious. That fundamental decision shaped everything. "Today's Puzzle" became "Today's Discovery," and "Completed" became "Discovery Unlocked." Small changes that completely shifted the emotional tone of the experience.

The sound.
This was the most persistent struggle. The first audio AI suggested felt like an alarm. When I asked for relaxation, it swung so far the other way it felt like a meditation app for people already half asleep. Then when I asked it to pick a track, it listed sources and told me to choose one myself!?

AI could describe what I needed. It couldn't make the call about what felt right. So instead of forcing one interpretation of "calm," I then added an audio toggle.

Failing to solve the problem produced the better design decision. AI could remix patterns. It couldn't bring the experiences and emotional connections that give those patterns meaning.

What surprised me

Once, without prompting, AI suggested numbering puzzle tiles to support players with visual impairments. I hadn't considered it, and it made me rethink an area I'd unintentionally deprioritised.

But the feature wasn't tested or validated, so I couldn't claim it improved accessibility just because AI suggested it. AI can broaden your thinking. It can also create a false sense that something important has been handled.

Where I almost lost my own point of view

The most uncomfortable moment wasn't where AI failed. It was where I stopped questioning it.

Some AI-generated screens looked polished enough that I accepted them without much thought, not because they were excellent but because they looked familiar, and my critical thinking subconciously started to switch off. Only later when I deliberately changed features such as the horizontal progression, blurred undiscovered landmarks and simplified the visual hierarchy did I realise I'd quietly moved from designing to just curating.

The danger isn't that AI produces bad work. It's that it produces work that's good enough to stop us asking whether it's the right answer and challenging AI and ourselves. As a design leader, I think that's one of the biggest risks we'll face moving forward.

What this means for leading a team

This was a solo experiment, but it changed how I'd lead a team through AI. Here are some things I would look to implement:

Set pinciples before prompts.
Without clear design principles upfront, AI naturally fills the gaps with statistically likely answers. Fast becomes familiar, and familiar becomes average.

‍Build AI fluency, not dependency.
Teams are already using AI. The job is helping them see where it creates leverage, where it adds risk and where human judgement is still essential.

‍Review designs for intent, not just execution or quality.
Add a checkpoint in projects that asks "does this still serve what we set out to create?" rather than "does this look polished/good?"

Treat trust as a design responsibility.
As AI-generated content becomes more convincing, people will question what is authentic. Designers and design teams will need to play a critical role in protecting trust, transparency and confidence, and also creating relevant experiences.

Moving forward design leaders become guardians of judgement: knowing what should be built, what should be challenged and what should be protected.

Final reflection

When I started this experiment, I thought I was exploring AI. Looking back, I was really exploring myself.

The game isn't finished. Retention wasn't as strong as I'd hoped, the soundtrack still isn't right, and some ideas didn't survive contact with real users. But that's why the experiment was worth doing.

It reminded that it can't replace the experiences, judgement and curiosity we have that shape meaningful and relevant  products. The future of design isn't humans versus AI. It's designers who know how to work with AI while protecting what makes great design human.

‍Beta status: 5 of 20 invited users played over 7 days, it definitely needs to add in more hooks for return usage. Happy to add more testers.‍

PLAY TILETOPIA >