Designing for Cognitive Stewardship
Overview
In early 2026, I noticed a pattern: people who were heavy AI users were also reporting cognitive fatigue and burnout. AI turns work from an act of production to an act of judgement, asking users to approve this or review that. Multiply that hundreds of times across a workday and it becomes incredibly taxing, reducing a person’s ability to think clearly and make quality decisions.
I led a 3-month innovation workstream with a nimble team of of 4 designers and 2 engineers focusing on how we can design model behavior UX to mitigate this issue and support users’ cognitive functions instead of draining them. In the end, we built three customer-evaluated prototypes and a model behavior for adaptive decision callibration, learning about the user and surfacing what’s most important while silently handling the rest.
Approach
Desk research to ideation
I started by building a practical research primer for my team, pulling current product research, industry signal, and classic academic literature like Tversky and Kahneman’s cognitive heuristic work. Then I gave everyone homework: go use our own products with this in mind, come back with what you find and back-of-napkin fidelity solution concepts.
Three exploration areas
I held a workshop where we collectively syntehsized our insights and ideas into three focus areas:
- Reducing the amount of decisions a user needs to make by surfacing only the important ones.
- Helping a users at critical junctions by framing downstream consequences and allowing parallel exploration.
- And creating a shared human-and-machine-editable context document to ensure all agents are aligned.
I assigned ownership of each focus area aligned with individual interest and ideation, and we moved into the next phase of rapid prototyping.
Rapid prototyping and iterating
The team had freshly gotten access to Claude Code and was eager to get building, but I quickly noticed that without up-front intent, the designers were going deep in areas tangential to the work, some going as far as building a small OS.
I felt it was critical to slow the work down a bit to make sure everyone was moving in the right direction. I built a ‘model behavior specification’ - a short, written document I asked everyone to fill out for their work and review collectively. It included the problem, hypotheses, a user story, and description of the ux. By pausing, even for a few days, we were able to go back into building coded prototypes that were much more aligned to the work.
As the work moved forward I was adamant about gathering feedback early and often. On top of team crits multiple times per week, I gathered a handful of concept testers on a bi-weekly basis to walk through the prototypes, turning the feedback around the next day for the team to move forward with. This proved to be incredibly important to make
Decision callibration behavior: surfacing what is most important
I ended up stepping in to lead the decision callibration workstream, designing the UX and the model behavior. I built a research-backed rubric the model would use to identify and evaluate if a decision should be surfaced to the user or handled silently and a series of prototypes to explore interaction patterns.

Human and Machine Evals
To refine the model behavior, I had over 70 people rate a battery of workplace decisions against the rubric, then evaluated the same set of decisions with 5 frontier LLMs. This allowed me to see the differences between people and models, and helped me refine the rubric and system prompt to better match human interpretations. It also highlighted the need to calibrate to the person over time, as certain preferences overwhelmingly influence delegation expectations.
UX Prototypes
Now with the refined behavior in hand, the question became how to surface it in a meaningful way.

Me and the team explored a number of options and landed on a combination of doorbell questions when the decision would determine direction and inline questions when the decision is significant, but does not materially change the scope of work.
Outcomes
In the end, the team had built 3 research-backed prototypes for their concepts and the refined, researched model behavior for decision callibration. Our outputs have been handed off to two product teams to begin A/B testing in production and to an applied science team to further refine and research the effects.