AI Pulse & Data Waves

GenAI for Good Is Being Built Backwards

Every major tech conference now has a "GenAI for Good" track. Every foundation has a grant program for it. Every AI lab has a responsible AI division with a slide deck about it. The label has proliferated faster than…

Every major tech conference now has a “GenAI for Good” track. Every foundation has a grant program for it. Every AI lab has a responsible AI division with a slide deck about it. The label has proliferated faster than any working definition of what it actually means to design one well.

The intent is genuine. The urgency is real. But when you look closely at how most GenAI for Good projects get built, you find a consistent structural failure: they start from what the technology can do, then search for a social problem to apply it to. This is backwards. And no amount of ethical review, fairness audit, or responsible AI checklist bolted on at the end will fix a project that was designed in the wrong direction from day one.

The design failure in GenAI for Good is not a values problem. It is a sequencing problem. Impact should define the question. GenAI should be the answer only if — after honest evaluation — it is actually the right one.

“GenAI for Good” is currently doing the same work that “AI for Good” did before it, and “Tech for Good” did before that. Each wave inherits the same fundamental confusion: the belief that attaching a powerful technology to a social cause is, by itself, a design process.

What’s different about GenAI specifically is the scale of the gap between what the technology appears capable of and what it can reliably deliver in high-stakes, under-resourced, and structurally complex social contexts. Language models are extraordinarily good at producing fluent, confident, and contextually plausible outputs. They are not extraordinarily good at producing correct, equitable, or trustworthy outputs in domains where the cost of error is borne by people who had no say in the system’s design.

That asymmetry matters. A hallucinating model deployed in a commercial chatbot is an embarrassment. A hallucinating model deployed in a healthcare triage tool for a community with limited access to alternative care is something else entirely.

The label does not come with the framework. Teams have to build that framework themselves — and most aren’t.

Here is how most GenAI for Good projects actually begin. A team — well-intentioned, technically capable — looks at what the current generation of models can do. They can generate text at scale, reason across documents, translate, summarize, and produce structured outputs from unstructured inputs. These are genuinely remarkable capabilities.

When you start from capability, capability becomes the ceiling of your ambition. You end up designing for problems that fit the technology, rather than designing for problems that matter. The communities you are ostensibly serving become audiences for a product demonstration rather than co-designers of a solution. The “good” in GenAI for Good becomes whatever the model can plausibly address — not whatever the actual need requires.

If the model can summarize legal documents and you’re building for underserved communities with limited legal access, the question is not “can GenAI help here?” The question is: what does this community actually need to navigate the legal system, and is a summarization tool the highest-leverage intervention available? Those are different questions. Most teams only ask the first one.

This is not a technical failure. It is a design failure. And it happens before the first line of code is written.

Most GenAI for Good projects share three structural design failures that distinguish them from commercial AI — not in obvious ways, but in ways that quietly undermine impact at scale.

The stakeholder is not the user. In commercial AI, “user” and “stakeholder” are roughly equivalent. The person using the product is the person whose behavior you’re optimizing for. In GenAI for Good, this is almost never true. The person interacting with the system is often not the primary stakeholder — the community, the institution, the ecosystem they belong to is. Designing for individual user experience in these contexts frequently means optimizing for something that doesn’t actually improve collective outcomes. A GenAI-powered mental health tool that feels helpful to individual users can still systematically fail a community if it displaces professional care infrastructure or creates dependency without support structures.

Success is not engagement. Commercial AI is measured by retention, conversion, and engagement. These metrics are seductive precisely because they’re legible. But for GenAI for Good, the most successful possible outcome might be that people use the system less over time because their underlying situation has improved. Harm reduction is not the same as usage reduction. Equity of access is not captured in daily active users. Teams that inherit commercial success metrics for social impact projects end up optimizing for the wrong thing — and the data will tell them they’re succeeding when they’re not.

The goal is exit, not growth. This is the sharpest difference and the one most teams resist acknowledging. Commercial AI is built to scale, to grow, to become indispensable. A well-designed GenAI for Good project should be built to become unnecessary. If the intervention works, the conditions that made it necessary should change. That requires designing for handoff — to communities, to institutions, to infrastructure that outlasts the technology. Projects that don’t plan for exit don’t plan for success. They plan for dependency.

The alternative is not more ethics reviews or more diverse datasets, though both matter. The alternative is a fundamentally different sequencing of design decisions.

Impact-first design begins with a theory of change, not a technology stack. Before any conversation about what GenAI can do, the team must answer a harder set of questions: what does success actually look like for the people this is meant to serve? What are the structural conditions that produce the problem? What would need to change for those conditions to shift? Who has tried to address this before, what did they learn, and why didn’t it scale?

These questions are not AI questions. They are social design questions. And answering them rigorously will frequently reveal that GenAI is not the most important intervention available — which is essential information to have before you build anything.

When the answer does point toward GenAI as a useful component of a larger solution, the design work changes character entirely. You are no longer asking “what can this technology do?” You are asking “what role should this technology play in a system designed around human outcomes?” That reframe changes every subsequent decision — what data you collect, how you evaluate the system, who you involve in the design process, and what success looks like.

Define the impact: what does meaningful change look like for this community? Map the theory of change: what causal chain leads from intervention to outcome? Identify the constraints: what do existing approaches miss, and why? Evaluate GenAI’s role honestly: is it a fit, a partial fit, or a distraction? Design for exit: how does this project transfer capability rather than create dependency? Build in accountability to stakeholders, not users: who does this system actually answer to?

The practical difference between capability-first and impact-first design shows up in a specific set of decisions that teams make early and rarely revisit.

It shows up in who is in the room when the problem is defined. Capability-first teams define the problem among people who understand the technology. Impact-first teams define the problem among people who understand the context — ideally, including those closest to the conditions being addressed.

It shows up in how success is measured. Capability-first projects track outputs: queries processed, documents summarized, hours saved. Impact-first projects track outcomes: did legal access improve? Did health outcomes shift? Did the community’s capacity to navigate the system increase? These are harder to measure, slower to emerge, and more honest about whether the project is actually working.

It shows up in how the team talks about GenAI internally. On capability-first projects, the technology is the center of the design conversation — its limitations set the boundaries of what’s considered possible. On impact-first projects, the technology is subordinate to the design conversation. When the model can’t do what the community needs, the answer is not to redesign the community need around what the model can do. The answer is to find a different approach.

That distinction sounds obvious when stated plainly. It is not obvious in practice. The technology is expensive, impressive, and heavily invested in. The social problem is complex, slow-moving, and resistant to measurement. The path of least resistance is always to let the technology define the scope of the possible. Resisting that path is the actual design challenge of GenAI for Good — and it starts long before the model is selected.