The Rise of the Orchestrator IC
On building AI agent teams, maintaining quality, and being the grown-up in the room — a conversation with Saeideh Bakhshi, Quantitative Researcher at OpenAI
I have been following Saeideh Bakhshi's writing on Substack for a while and am impressed by her earnest takes on AI UX research and her generous sharing of wisdom with the community—she even created a Research Toolbox, a custom GPT to help you improve your research projects. After reading her piece on wicked problems and AI, I decided to reach out to her, and we had a wonderful conversation.
Without further ado, enjoy!
Kursat: Thanks so much again for accepting my invitation. To start, could you briefly tell us about yourself and your journey as a UX researcher? What brought you to where you are today?
Saeideh: I’m a quantitative UX researcher, and I’ve been in the industry for 13 years, working across several companies — mostly in roles like UX researcher, quant UX researcher, or some version of research science. I come from a computational background; my PhD is in computer science. But early in my PhD, I became really interested in social computing, computational social science, and the HCI side of things, so I switched my topic to that. I’ve been researching at the intersection of complex technological systems and people — how humans interact with those systems.
And over the past 12, 13 years, most of my work — even before generative AI — has been on these more complex intelligence systems: ranking, recommendations, how people perceive them, how they build trust, how they work with them. My fascination lies in how we make technological systems accessible and useful to people.
Kursat: And you’ve been doing that mostly in tech companies?
Saeideh: Yeah, always in the tech sector. I started my industry career at Yahoo Labs — I was part of the HCI group there, working on how people engage with recommendations on photo-sharing, like Flickr and Tumblr. During my PhD, I studied how people engage with multimedia content on social media. Then I moved to Facebook, where I worked on video recommendations, and on some of the technical aspects of the experience, like performance and reliability, and how people perceive those across different cultures. Then I went to Google — most of my time there was spent on search and the assistant: how people interact with search results, what are the new ways people seek information, and how intelligent systems change the way they interact with technology. Then I spent a short time in fintech at Cash App and went back to Meta, focusing again on recommendation systems — personalized recommendations in the feed and on video across Meta products. And then I joined OpenAI last year, focused on quantitative UX research in general, and on understanding how people interact with models.
Kursat: What’s your typical setup on your team—do you work with designers?
Saeideh: I worked most closely with designers between 2016 and 2023, during my first stint at Facebook, and then at Google, where I was actually part of a design team inside an innovation group. But that was mostly because of the nature of the work; it was more product- and strategy-focused. After that, when I worked on recommendations at Meta, it was less so, because that’s more on the technical back end — so I partnered more with engineers and researchers.
Three modes: tool, collaborator, and something more agentic
Kursat: Great, let’s talk a little about how AI has changed the way you actually work, day to day — what part of your process has changed the most? You’ve been in this space for a while, so it’s not exactly new for you.
Saeideh: Yeah, I’ve actually been reflecting on that. My usage has changed a lot over the years. I’ve been a power user since ChatGPT launched in 2022, experimenting with it along the way. And I’d say there are two or three stages to it.
Initially, it started as: a little help with parts of what I do — on writing, on organizing a document, on outlining, on ideation, on generating new ideas. So, more like a tool.
Then there was a time I felt like I was using it so much as a second brain, or a cognitive scaffolding tool, where I’d go and say, “Here’s an idea I have — push back on me. What are the flaws in it? What would someone from this discipline take from it?” Just extending where I am, as a collaborator, would improve me. And during that time — maybe a year or two — I actually saw immense growth in my own skills as a researcher, as a thinker, as a person, because it was pushing me a lot, and I was pushing it to push me.
And then recently, with these agentic tools and more ability to actually do things, it’s shifting again — now it’s not just a collaborator but a way to get things done, somewhere I can delegate. It’s kind of like a smart intern: you give it a bunch of tasks, you give it feedback — “you did great, keep doing this” — until it’s independent enough to do that thing on its own. So that’s the conceptual arc.
Kursat: I love it. What I’m hearing is there are these different modes — one is more of a tool, then there’s the collaborator or thinking partner, and the third is almost like you have a small team — agent interns, junior teammates — and you delegate to them. Is that right?
Saeideh: Kind of like exactly that — now that you say it, it resonates. It started with a tool, then became a collaborator, and now it’s more like a system builder, a process collaborator, or a program manager kind of thing that can do things end-to-end. More like a very smart assistant.
What gets automated, and what she keeps
Kursat: Do you have a concrete example, especially for the agentic part?
Saeideh: Sure, on the research side, my process usually has two or three phases. There’s the problem definition and scoping — do we even want to solve this? — which involves a lot of stakeholder communication: “This is the problem we want to solve, and these are the approaches.” Then, once it’s scoped and well-defined, there’s data collection, insight generation, and analysis. And then the final stage: now you have the data and the findings, and you want to create interpretations you can communicate, a call to action for what we need to do in the product sense.
The part I’ve been automating most is that middle part. I do the scoping myself — I gather context, figure out the right way to do it, scope it — and then I pass it to the AI: “I have these particular research questions, I want to understand these user groups, here’s the data source, here’s the analysis I’m thinking, I’m open to additional analysis, these are the decisions I want to drive, and this is the final output I want.” That part is now basically fully automated. There are checkpoints where I ask, “show me what you’ve got,” and say, “no, I don’t agree with this, maybe try this other type of analysis.” That used to take weeks — you had to write code, or collect data you didn’t know where to find — and now it takes hours.
And then the final stage, the meaning-making — that’s the part I still try to do myself, or work with AI in a scaffolding way. Once I have the list of findings, I want to interpret them myself; I want to be more hands-on there.
Kursat: So you own the beginning and the end — you give the operational stuff to AI, and you keep the scoping and the final meaning-making. You touched on this in some of your recent writing: scoping is almost a wicked problem, and the operational part is a tame one. Can you speak to that — and to your definition of a wicked problem?
Saeideh: Yeah. An example: if I hire someone straight out of grad school into industry, they usually know the methods — they know how to run things rigorously and do analysis. But they usually can’t do an end-to-end project in a product or business space because they haven’t seen it, don’t have the context, and haven’t been in the rooms. And in my experience, that’s the part that takes years — it’s a kind of human experience that accumulates. It comes from social context, from working with different people in different places and situations. It’s not easy to formulate or operationalize.
So those are the things that require human judgment, and that’s what makes it a wicked problem. It’s hard to put on a playbook and say “this is the right way,” because you have to understand human motivation, institutional knowledge, and how different people work with it. You have to have past experience of which kinds of insights fly and which don’t. A lot of that is situated understanding — social understanding that AI doesn’t have. The final part, the meaning-making, is the same: not because AI can’t interpret results, but because it doesn’t have all the social and business context.
When I think about a wicked problem versus a tame problem, a tame problem is one that’s very easy to define, mostly that middle, operational part. The scoping is easy, the process is well-defined, and you get a specific result. The wicked problem is the one that requires all of the rest.
Kursat: Right — you need to do a lot of reading of the room, all that contextual, tacit knowledge. The intuition you grow over the years that a new grad doesn’t have yet.
Saeideh: Totally.
Where the practice moves: upstream, and the governance problem
Kursat: When it comes to AI changing UX design and research — do you see more change happening downstream, in execution, or upstream, in strategy?
Saeideh: I have some not-very-well-structured thoughts here. One is that it’s much easier to generate insights now. A product manager can just ask AI, “What do we know about this?” and get a summary, some data, and some visualization. So that part has become much easier, and everyone can do it, which means it’s less of a differentiator for an insight professional to generate insights.
But on the other hand, there’s so much nuance — tacit knowledge, contextual knowledge, methodological knowledge — that goes into doing this right. And the humans who are decision-makers are being flooded with insights, but they don’t have the attention span to consume them or make sense of them, and they don’t necessarily trust them all because each one may tell a different story. So the way I see researchers going upstream is: are you trusted enough and experienced enough in your craft to identify what’s important and make meaning across all these sources? The person making a decision might only have one hour to see this, so it becomes a synthesis on top of a synthesis, meaning on top of the meanings.
The other thing is rigor and craft — are you differentiated enough? If you’re not deep on the skill, you can’t identify the problem. I’ve seen so many analyses go wrong because they didn’t pay attention to the details. You can ask for an analysis, but if you don’t have the context of how to handle missing data, you can get completely different results. Someone who isn’t data-literate won’t know that. So you have to have that knowledge to ask the right questions, to design research systems that get at the right thing, to become insight system-builders who make sure the quality is there.
We see the same in engineering: everyone writes code, but is the code maintainable? Does it break when it’s part of a bigger system? Someone has to monitor it; someone has to know where it went wrong — and that knowledge is usually held only by experts. It’s happening in design too: everyone can make an app now, but people don’t have the attention for hundreds of apps. Do we need apps for everything? What’s the right thing? How do you get the signal out of the noise?
Kursat: In design, there’s so much emphasis on taste, on curation — finding the right design, the attention to detail and quality. On the research side, it sounds similar: access to analysis is so cheap now, the way creating UIs got cheap and fast, but the quality still isn’t there. It still depends on the experts owning it.
Saeideh: We hear a lot about taste and judgment, but I think even that can be broken down — what do we mean by judgment, by taste? It could be that the quality isn’t accessible to these systems because of the intuition and human experience we bring. And it might be more at the system level than the task level. So, in general, how do we want to define what taste means? It’s time to define these concepts in more concrete ways.
And it’s up to us practitioners to define our value and approach beyond AI, and we have the agency to do so. And there’s going to be more need for some sort of governance over the output, because if everyone can create a prototype or an insight, who decides what’s good? That’s where the taste, the expertise, the quality comes in.
Kursat: By “governance,” do you mean the guardrails and how you define quality?
Saeideh: Yeah — define quality, and extract the high-quality thing out of the pool of everything being generated. At the end of the day, these systems are for humans, so someone has to surface what’s actually worth paying attention to. And when you generate a lot, there are downstream problems too — if you generate 50,000 lines of code, it’s harder to know where something went wrong. It’s the same with insights: a lot gets generated, and when a decision needs to be made, how do you know which one is actually relevant and reliable? So that post-generation step — once you’ve generated the content, the code, the prototype — what governance ensures it’s high-quality and won’t break the system or the decision?
Kursat: Can you give an example of that governance?
Saeideh: One example: I come from a computational background, so I’m skilled with stats and data analysis. But many people now do quantitative analysis without that background, and they can still produce very nice-looking insights. And it might be only me who can look at it and say, “What about this?” or “this statistical test doesn’t make sense here because it’s not the right metric.” So I kind of have to be the grown-up — the reviewer — for a lot of this experimental work, and build some of those policies to make sure things are right.
The other approach is this: you want to enable people to do this, but they don’t have the right system in place. So, based on your expertise, how can you build systems that democratize that access?
For example, I built a custom GPT to do survey reviews. The way it works is that I went in and put all my knowledge — what makes a good survey and the potential problems people don’t catch — into the custom prompt. Then I made it available to my team as a first step. They’re going to use AI anyway, so at least put that knowledge into the first step: a custom GPT that flags issues and reviews them for them. Once they’ve done that, they’re in a good place, and then I can do the manual review as a final check.
Kursat: So you have two lines of defense for quality — AI does the first round, and you do the second.
Saeideh: Yeah, I just do the quality call, because at this speed it’s not possible otherwise. Everything’s moving really fast, everyone’s empowered and enabled to do things — you don’t want to slow them down or be the gatekeeper, but you also want to keep the quality bar high.
AI as a teammate — and the accountability boundary
Kursat: I wanted to touch on something from my own work — I’ve been exploring this idea of AI as a teammate, and how we can intentionally design AI teammates, like through a custom GPT, crafting a system prompt to almost create a teammate. Curious about your take, and where the boundaries of that framing are.
Saeideh: On a personal level, I totally see AI as a collaborator, an enabler, something I can rely on to make things faster, more efficient, or to help me think. But the teammate framing — I’d challenge that, because there’s no accountability with it. When I think about a teammate, I’m not just thinking about delegation or collaboration; I’m thinking about accountability — they’d be responsible for what they produce. With AI, we can’t hold it accountable. We’re using it, and we are accountable for the output.
So that’s mostly the sense in which I hesitate. I described some of these uses — thinking partner, helping generate a draft, pulling things from Slack, summarizing, analyzing, and writing code. I don’t write any code anymore, just natural language; I’ve totally lost touch with that level, so there’s that abstraction. But I still don’t call it a teammate, because that’s not my mental model of a team.
Kursat: What I’m hearing is that the boundary is accountability — you can’t really expect it to be accountable.
Saeideh: Right. On the other hand, I’ve been an IC for most of my career, and I feel extremely empowered now because I can run multiple agents, create agents, and assign different tasks to different agents. In the past, as an IC, you had to get things done through influence — building relationships to get people to do things for you, because you didn’t have management authority. This is really good for me in that sense; I can delegate.
But again, because it lacks accountability and I feel responsible for anything it generates, the bottleneck becomes me. I have to review it. I have to feel comfortable putting my name on the output. So it’s not like being handed an empowered team of smart people — that’s the deal.
Kursat: That insight about AI agents empowering you as an IC is excellent. It brings to mind Stanford research indicating that individual contributors are becoming managers of AI agents. In that sense, they need product management skills to oversee them effectively.
Saeideh: Yeah, exactly — you need some management experience.

Kursat: So, as one AI gets more capable and acts more like a teammate, but still without accountability — how do you see that playing out?
Saeideh: Let me give you an example of a research approach where I use AI. There’s this concept where you get qualitative feedback from users and want to infer themes from it. The rigorous way to do this is to have two or three researchers, each sitting in a separate room, reviewing the data on their own, without talking to each other — going through all the feedback manually, tagging things (”this is a pain point,” “this is an interface issue”), and building a taxonomy of tags. Then those researchers sit together and discuss what they agree on and disagree on; the agreements become common themes, and the disagreements are reasoned through. That’s the rigorous way.
But in industry, you usually don’t have that much time or that many independent people. So I use independent agents. I created four different agents, each with a role, each reviewing the data independently — they don’t share context; each has their own spreadsheet and deliverable. Then I have another independent agent review those and decide the agreement and disagreement. Then the agents talk to each other, and at that stage they show me what all the discussion came up with — and I put my input in there too. So there’s a human input at the end. And honestly, that’s more than you could realistically do with a human team. You can even prompt each agent to take a different point of view — “you’re an economist, you have this kind of view,” “you’re a psychologist, you have this one.” The opportunities are huge.
Kursat: Do they all use the same model, or can you set different models too?
Saeideh: I don’t know how much you’ve experimented with different prompts and settings, but the prompt, the context, and the added constraints make a huge difference — they act very differently even on the same model.
Kursat: Yes, this is a great example. In a way, you created your own research team. I love that.
Saeideh: Yeah, make your own team. I’ve also been thinking a lot about creating persona-based agents — like a designer persona who’s very vision-forward, or another who’s very focused on aesthetics — and then asking what they’d think about my artifact. There’s so much opportunity. But it’s up to us to experiment and figure it out.
Reflections: surprises, a change of mind, and advice
Kursat: A few more reflective questions. Where has AI surprised you the most — positively, or negatively?
Saeideh: In a positive way, looking back at the past couple of years, how much I’ve grown because of AI is just amazing. I’m also neurodivergent, so traditional learning and communication have always been a little hard for me — the way I learn and organize things is very different. Having ChatGPT personalize to me — “I’m curious about this concept, show it to me visually,” or “create a high-level framework,” or “I need the end goal before I can make progress, so what are some potential angles” — that’s been amazing for brainstorming, articulation, and giving more organization to my own thoughts. That’s helped me a lot.
In a not-so-good way, one thing that sometimes worries me is the first-pass priming effect. When you wanted to think about a new problem space, in the past, it was a blank page: you put your own ideas down, you had to come up with some of that friction, and you had to form your own ideas before you could ask for help. That first step is really important for forming a point of view and coming into a problem as an individual. Going to AI too quickly makes it much easier to get to that first step — but then I worry, am I getting too biased toward the frame it gave me, versus my own framing? There’s a psychology thing about the priming effect — starting from an existing state rather than a blank one biases you toward it. So I’ve been trying to exercise the opposite: pen and paper, even rough bullet points of my own thoughts first, before I go to AI. It’s just my worry that I’ll lose that ability.
Kursat: Yeah — there’s that recent Nature paper about cognitive surrender, losing your thinking ability over time. But I like the priming-effect lens too; that’s a good way to think about it.
Saeideh: I feel like it’s still not a fault of AI — it’s something we’re accountable for. We need to make a conscious effort, and that’s hard.
Kursat: Right — you can just type it up, and that’s the tension: you have to resist that ease. It’s very easy, but it shouldn’t be.
Saeideh: Exactly. You’re the designer — you know that friction is sometimes good. It brings creativity.
Kursat: Yes — the struggle is part of how you grow.
Saeideh: Exactly.
Kursat: A couple more. You worked in this space before generative AI became a thing — what’s a belief about AI, design, or research that you’ve been trying to introduce more nuance into, instead of seeing it as black and white?
Saeideh: I was a very early researcher in social media — my first PhD research focused on Tumblr, Instagram, and Pinterest. And there’s always been this tension: is social media good or bad? There have definitely been negatives, but there’s also a lot of research showing how much it has improved people’s lives — creating social capital and bringing people closer together. So I think we’re at a similar intersection now with AI: there’s going to be both good and bad, and I don’t see it as black-and-white.
I think we have two different responsibilities. One is as a user — making an intentional effort to use it as a tool that provides cognitive scaffolding and helps you become a smarter person, rather than offloading your thinking in a way that makes you dumber. As a practitioner, how can I use it to become a better researcher, rather than just dumping all my work into it?
The other aspect is as a community — the way we frame this. You hear a lot of “let’s automate this, AI can do that, we won’t have jobs anymore.” I think those black-and-white narratives aren’t fair either. And we can try to change that, as people who have influence over how this technology gets designed. That’s what I’m excited about: how do we make this a genuinely useful thing for the next generation?
Kursat: Yes — holding that nuance, and recognizing the agency we have, especially as designers and researchers, to influence it for the better. Final question: What advice would you give researchers or designers working in this space as they navigate this AI moment?
Saeideh: It’s a very unexplored space — all of us are figuring it out — so just try things, adapt, be creative. But one thing I keep telling people: there’s so much new, so much distraction out there, so don’t try to do everything. A UX researcher asks me, “Should I do vibe coding? Should I learn to code?” No — you do it if it helps you be a better researcher, if it helps you have more impact, if it makes your workflow faster. Ground yourself in your mission — in your life, your industry, your job — and then use AI as much as you can to improve that, rather than chasing all the noise. I’m a researcher; I don’t need to build ten different apps. But if there’s an app I can build that makes my insights more actionable, that’s worth it.
Kursat: Back to the intentionality piece. This was really fun, Saeideh—I appreciated your deep thinking and generosity in sharing your wisdom today. Thank you so much.
Saeideh: Thank you so much. It was fun. I look forward to seeing what comes out of it.



The power shift you're describing is real, but it creates a fragility that doesn't get enough attention.. When individual contributors are running agent fleets, the organizational load-bearing capacity shifts to whoever's willing to review and certify outputs.. That's not always the most senior or experienced person -it's whoever has the bandwidth, the judgment, and the appetite to absorb the accountability that comes with signing off. Organizations that don't architect for this are going to end up with accountability sitting with whoever happened to be the last human in the loop, not whoever should actually own the decision.
"a program manager kind of thing that can do things end-to-end. More like a very smart assistant." contradictory...