AI as a Senior Hire, Not an Intern
Jess Holbrook on Managing AI like a Senior team member
A year ago, I met Jess Holbrook at the Swell Conference in Honolulu, where he laid out a set of “primitives” for designing with AI as a material: chat with everything, remix, semantic resize, multiplayer, and attention. I wrote it up at the time (The Emerging Shapes, Makings, & Principles of AI Design), and the framing stayed with me all year: a shared vocabulary for a material that’s probabilistic, alive, and a little hard to hold.
Recently, I went back to him. Jess currently leads user research in Microsoft’s AI division — before that, he was the first UX researcher in Google’s machine intelligence group, started PAIR with Martin Wattenberg and Fernando Viégas, and led responsible and GenAI UXR at Meta. I wanted to know which of the primitives held up a year later, and what his team has actually been doing in the fifteen months they’ve spent remaking how research works.
Where he landed at the end surprised me. His advice for navigating the agent flood came down to a single word — sensitivity — and the case he made for it quietly reframed the whole “have taste” conversation.
Without further ado, here’s our conversation.
Enjoy!
Kursat: To start, can you tell us about yourself and your journey? You’ve been at Google, Meta. I’m curious about your origin story.
Jess: In the beginning was the Big Bang (laughs). Well, my background is originally in social psychology — that’s what I went to school for. During grad school, I picked up human-computer interaction after randomly working at a startup one summer. Not a cool startup — this was back when it was really ugly grunt work. And as an intern, you do a different thing every day, so one day they said, “Hey, we need you to design the page where our ads get sold. Here’s a copy of Dreamweaver. Go do it.” And I said, “Cool — what’s Dreamweaver?” This is dating me a bit, but Google wasn’t that mature yet, so I just started commenting out code. I’d go, huh, when I get rid of that code, all the images in the center move to the left. So I guess that code centers things. I did that for a while, went back to my advisor, and said — I didn’t even know what design was — I like psychology, I like technology, and I like art, because I didn’t know what to call ‘design’, so I just called it art. And he said, “Oh, you should be a UX researcher; go talk to this professor - Sarah Douglas.” She took me under her wing, and I ended up with a minor in HCI.
Then I graduated and — for people struggling in the market now — I applied to about 50 places. I had a spreadsheet. I got 49 nos and one yes: a contractor role at Microsoft, in the hardware division. I got an FTE position there and worked on Windows and Internet Explorer for a couple of releases. Then Amazon, on literally Amazon.com — two areas, international (why are we doing so poorly in Japan and China?), and design systems. To my knowledge, we did some of the earliest research on what makes a design system work and whether it’s even worth building one.
Then I joined Google in the cloud division — randomly, that’s what they were hiring for. I did their hello-world starter project before my interview and thought, well, this is terrible, so I’ll have a job for a while, because they really need to make this a lot better. About a year and a half in, came the most consequential move: I joined Google Research, the Machine Intelligence group. To my knowledge, I was the first UXer in that group, about 300 people. Being very early to the deep-learning wave proved a good call, looking back. At the time, it seemed terrible. People were like, “What are you doing? Why are you leaving Cloud?” Cloud was exploding. Everyone who wanted a big career was staying in the cloud. And I left, and they were like — what UX work is even there to do over there?
I worked with early computer vision models and some of the earliest language models, not LLMs. And the big thing there: I started a group called PAIR — People + AI Research — with Martin Wattenberg and Fernando Viégas, and ran that for a few years. Then I ran the responsible AI UX team for a bit, did the same at Meta — led the AI UXR team there for a year — and I’ve been at Microsoft for two years now, leading user research in the Microsoft AI division.
Kursat: You’ve been in all the highlights of the industry.
Jess: That wasn’t the plan, by any means. But it’s been great. It’s a total cliché, but it’s true: at these giant companies, you get to meet exceptional people. They’re aggregators of amazing people — there’s no other way you’d run into them in your life. Not to be cliché, but that’s the greatest gift of all.
“Research on AI, and research with AI”
Kursat: One of my earlier guests, Leonardo Giusti, made a distinction between designing with AI, designing for AI, and designing the AI. How do you frame your team’s work?
Jess: I only split it into two — maybe it should be three. I think of it as research on AI and research with AI. “On” covers product behavior and model behavior. “With” is how we thoughtfully integrate it into research practice to improve it.
Kursat: It’s simpler.
Jess: Always trying to simplify.
Kursat: Now, I want to talk about AI’s influence on your work. How is AI shaping your team’s practice?
Jess: We’re in maybe the toddler age of our AI transformation — probably fifteen months in. When I joined, we were using basically zero AI in the process. And I said, as the team leads, I don’t know the right number, but I’m pretty sure it’s not zero.
So we started with data collection — AI-moderated interviews, leveraging vendors like Outset.ai. And we checked ourselves by rerunning a study we’d already run, since we had a baseline. A control. Do we get better results, the same, or worse? Are they faster? Does it actually take more time because you’re learning a new tool? We found that we got comparable results faster.
It’s interesting to me because the conversation is all about synthesis and analysis, and people forget that it was really bad until, honestly, a couple of months ago. That’s one of the challenges: if you’re not checking the tools at least monthly, you’re probably already behind. I had a note on my Substack the other day that got like five hearts — the last five months have been the weirdest five years of my career. People are feeling that.
Kursat: Yes.
Jess: Then, alongside that, my design counterpart, Liz Danzico, and I ran an upskilling effort across the whole studio — designers, researchers, writers, and UX engineers. We call it The Royal Academy of AI. Very silly, absurdist on purpose, because that creates a safe, participatory space for people to explore without worrying about looking dumb. Unblocking people on the simplest questions. When you start using VS Code, what’s a project? It’s okay. A project is a folder. That’s all it is. Now you’re unblocked.
On the research team, we’ve been building tools for ourselves. For example, the fraud detector. We started seeing participant fraud — people trying to get into studies. We told the vendor, and they said, “We’ll put it on the list; we’ll get to it.” So a researcher on our team just vibe-coded a fraud detector. That was a big aha moment for us because we realized that if we need to fix something and someone can’t fix it for us, we don’t have to wait.
The way I set up the team to experiment with and integrate AI into our process is bottom-up. My job as the lead is to create the conditions — a gentle pressure and expectation that everyone’s experimenting. I had a saying early on in our transformation: usage isn’t mandatory; experimentation is. I’m not forcing anyone to use any of this. I do expect you to try it. And if you try it and it doesn’t help, nobody’s making you use it. But every single person on the team has found something they use it for.
Then someone creates something, and it explodes. One person basically “site-ified” our reports — we don’t ship a lot of reports anymore; it’s a site you go to that tells the story, with data you can interact with and export. That blew up. Though there’s a novelty trap there — ooh, cool site — and you have to ask: is it actually helping us figure out what to do, or does it just look cool? Recently, we gave everyone the ability to build skills for the agents, and we ran a company-wide UXR skills showcase bonanza. Now there’s a common repo of skills anyone in the company can use.
And the thing we’re putting a lot of energy behind now: we have a common research repo — we use Marvin, plus Outset.ai, Listen Labs, Dscout, and Sprig for collection — and we’re building our own homegrown end-to-end research assistant. We don’t have a catchy name for it yet — it’s called “research collaborator.” It has the context of all our research, prompting, and contributions, and is iterated on by the whole team, with awareness of every skill people have built and access to design skills for communicating ideas.
Because here’s the cycle you hit: everyone experiments, all this stuff happens, and then nobody’s actually using anybody’s stuff. At some point, as a team, you have to say — okay, we see the options set; now we actually have to put energy behind something to improve it. This is when my role as a lead becomes more top-down. Nothing gets into an iteration loop on its own.
The discontinuity
Kursat: Where did AI surprise you?
Jess: The coding inflection was wild. At Thanksgiving, the coding models didn’t work. In December, they did. There was this abrupt discontinuity.
I remember my first “oh shit” moment. People kept telling me, “You’ve got to try it.” So I had it build a site — it’s called Plusmaxgoone. It’s just every product that uses the words “plus,” “max,” “go,” or “one” in its name, because I find the naming so unimaginative it cracks me up. It runs weekly and updates itself. I one-shotted it, and it came up in thirty seconds, and I just sat there and did the out-loud “holy shit.”
And then the researchers are building tools for themselves. There’s always a tool you need, always a new process. We do a lot of evals, so some folks on our team built a pipeline that makes the eval data clear — and the examples behind the data clear. You can see exactly where we win against ChatGPT, Claude, or Gemini and where we don’t, and it all exports cleanly into the pipelines. That’s when it felt like, okay, we’re building systems now.
I’ve been building little utilities for myself, too. On a Mac, when you minimize windows into the dock, they get all mixed up. So I built a little utility — group all my Ghostty windows, group all my Slacks, group all my Teams. Just a thing I want, so now I can make it.
Kursat: Claude Code or Codex?
Jess: That one was Codex. I use both — trying to stay aware of all of them. I think people have slept on Codex. I had a couple-day fling with Fable as a lot of people did, and that was amazing, but when Fable went down, I went back to Codex, which I’d found robotic a few months ago — and now it’s really good. Probably my main now.
The teammate trap
Kursat: What’s your POV on AI as a teammate? Where’s the boundary of that framing?
Jess: I’m on record from years ago as being very anti-anthropomorphic language. And I’ve mostly, almost entirely changed my mind about that. Most of it isn’t as harmful as people think — it’s a shortcut, a way for humans to tell other humans what’s happening. If you and I are at an ATM and it’s taking a while, and I say “oh, it’s thinking,” neither of us is confused into believing the ATM has a mental state. If I say “Codex thought the code was in this repo when it was in that one,” I don’t think it has a mind — it’s just a far more elegant way to talk than “it mistook the reference.” So I’ll talk about it anthropomorphically, but to be clear, I don’t think there’s an inner life there.
The biggest thing I see with treating AI as a teammate is that people fail when they try to be too smart with it. A lot of the time, you want to give it space. I’m really into the self-improvement loops now — I have a skill that runs each week and says, reflect on everything we’ve done and create new skills based on what you’ve seen. I’m not prescribing anything. Just — look at everything, tell me what you see, go do something about it. That tends to get you more creativity. When you really need it nailed down, you give it structure. But I often find out what I actually want through that loop.
I don’t think of it as a full-blown teammate, and a big part of that is the jagged frontier. I still get blown away — this is amazing; it’s doing so well — and then it’ll make the weirdest error. You realize, oh, I fell into treating it fully like a person, and you get betrayed by that. So I think it’s this weird third thing. People love the tool-versus-teammate dichotomy, but the most value I get is somewhere not dead center — probably closer to teammate, but without falling for the teammate trap, because that bites you.
One mental model I like: it’s an alien intelligence that studied humans for a long time, which is sort of what pre-training is. But it still messes things up. It’s kind of like Alf. Super smart, observed humans a bunch, interacts with us really well — and then it tries to eat a cat every once in a while, and we’re all going, no, no, no, don’t do that.
Kursat: I love that. When you write instructions, do you have a framework? I’ve been borrowing from team design literature — Hackman: clear goals, structure, and resources.
Jess: For me personally, I let the AI take the wheel. I’ll say I want a project about blank, and if I have opinions on a couple of things, I’ll engage in dialogue with the model to specify them. For example, I have a paper analyzer — I give it academic papers, and through iteration, I found a format I like: what the paper is about, the major claims, the evidence, and the strength of that evidence. I didn’t handwrite much of it. A lot of times, the models come up with a way to express my intent through an instruction set better than what I’d have written. I also told it never to dead-end — for every paper, suggest others I should read that I haven’t seen. So I’m imbuing it with intentions, not directives. It’s more like managing a senior person than a junior one.
Which is why I find the framing of “AI is like an army of interns” so funny — that’s my best-performing Substack post of all time.
Never once in my entire work life have I thought, “You know what we need right now? An army of interns. Hundreds of people who don’t really know what they’re doing. That would really help.” It was a total Rorschach — half the comments were “yeah, AI sucks,” and I had to come back and say, I don’t hate AI; I just think it’s a terrible metaphor. I get more out of it by treating it like a senior researcher rather than a junior one.
Kursat: I’ve heard that intern metaphor so many times.
Jess: The one I really hate is “we’re all builders now.” I gave a talk literally titled “We’re All Specialists Now.” To be clear, we all can build now, which is awesome — I’m probably doing too much of it. But the idea that we all become this generic builder goo in the middle? You know who that sounds good to? People who employ people and want them interchangeable, so they don’t have to pay for specialization. That’s just what specialization of trade is — someone creates a specialty, captures value, it gets commoditized, it becomes part of the stack, and people go on to specialize in the next thing. The “we’re all builders” narrative assumes humans spawned from the mud without context. Like, have you met a human? You do realize that studying something for 15 years has value, and you shouldn’t toss it because you can code now, right?
But people should build. Making things, caring for things, maintaining things — it’s good for the soul. When someone’s in a really bad spot, burnt out, I ask two questions: what’s the last thing you made, and who’s the last person you helped? Go pull some weeds, do the little project you’ve been meaning to do, go help somebody. You’ll feel better.
It’s the same reason I tell team members not to be available on email or Teams or whatever when they’re on vacation if they offer. What makes you interesting as a researcher and a teammate is what happens to you outside of work. I don’t think of you as an email-response machine. Go live your life so you can keep being this unique person who adds to the team.
The primitives, one year later
Kursat: At Swell, you gave us the primitives — chat, remix, these new materials for design. Has anything changed?
Jess: Chat with everything — yes and no. Chat is everywhere, but it isn’t really imbued in many different places; it’s still centralized. I still go to the one Claude, the one ChatGPT, the one Copilot. Anthropic’s integration announcement is closer to chat-with-everything
Semantic resize went mainstream. Text is infinitely malleable now — every word processor has the little panel: want me to summarize this?, want me to expand it? What I haven’t seen enough of is using that for accessibility — e.g., we’ve done work on modifying chatbot outputs to be more useful to neurodivergent users. There are lots of easy wins there that just need to be prioritized.
Remix — still mostly contained in the Midjourney Discords, but Nano Banana broke style transfer open, and the new MAI image models are quite good. The OpenAI Ghibli moment was the big remix moment over the past year.
Format translation — still not seeing as much as I’d like, and I think it’s massive for education. Oboe is my favorite here; you can see real care in the product, and I have a lot of love for people who have care. But I still don’t have the thing where I’m reading something and say, I’m getting in the car, turn it into a podcast — get home, turn it into a video. That doesn’t exist yet.
Review and supervision — completely mainstream. Write the doc for me. Oh no, now I have to edit someone else’s words (the model’s), and I’m not sure they did it right. More and more people are asking, “Am I even saving time, or did I just trade writing time for reviewing time?” We went from the tyranny of the blank page to the tyranny of reviewing the full page. We’re tormented either way.
Multiplayer — still zero. Ming-Li Chai and Rachele Benjamin on our team did some great work on social AI, and the finding is that people are hacking their way into making AI social — it’s just not native yet. AI is massively parallel play right now. I think social is the opening that cracks it into the really big consumer cases. A huge greenfield.
Attention — the agents. Exploded. It’s multitasking an order of magnitude beyond anything we’ve had. And it’s a casino — you’re hopping between instances, getting little productivity dopamine hits. I think Simon Willison said his brain is fried by 11 am from hopping between agents, and everyone knows that weird feeling: this is so cool, this is so cool, and it’s short-circuiting me at the same time.
Kursat: How does that get tamed?
Jess: We go up a level of abstraction — that’s the only way. We’re already starting with the/goal abilities and the meta-loops. I got into online chess over the last six months, playing multiple games at once, and it has the same effect — invigorating, then suddenly I’m going back multiple moves, asking, wait, what was I even trying to set up here? Rewarding and exhausting at once. So I think we keep adding layers of abstraction that work more on their own — (laughs) but then people will just spawn even more to fill the time. Everyone says technology will free up time for what we really want to do. Apparently, what we really want to do is scroll. Someone called agents “TikTok for productivity.” ‘Ooh, do I get something good? Did it finish the task on its own?’ It’s the intermittent variable rewards from the feeds all over again.
Kursat: I ask all our guests — any recommendations for researchers and designers navigating this wave?
Jess: There’s all the stuff everyone says, so I’ll try to be additive. I’ve been telling people they need to be more sensitive. That’s my version of “have taste” — and that whole taste conversation is broken too, mostly coming from people who don’t seem to have much and think it’s a static thing. Taste isn’t static; it’s evolving but, most importantly, relative to everything happening around you.
What I mean by sensitivity is sensitivity to people’s needs, to product quality, to your teammates. Because — to channel McLuhan — the medium is the message. What becomes of the message when the medium is agents and agentic workflows? They’re naturally distancing. They abstract. They put everything at arm’s length. And I worry that days filled with those abstracted interactions will be numbing, much like doom-scrolling. You look up after an hour and think, well, I got my dopamine hits, but I’m a little numb now.
One antidote is to stay very sensitive, in a good way: attuned to what is happening with this human, to how I can be of more service, and to how I can use this design as a greater expression of care for the people I’m designing for. I think that keeps us tethered to what matters and not dragged away in the flood of agents.
Kursat: That’s the care aspect — keeping the human quality, caring for who we’re serving. A beautiful place to land. Thank you so much.
Jess: Thanks for having me, and thanks for the conversation. I really enjoyed it.
Kursat: Me too. Have a great weekend, and let’s keep in touch.
Jess: Sounds good. Talk soon.



