How We Built Our AI Principles

The intentional process is a direct reflection of how we choose to approach emerging technology responsibly

Late in 2025, before we wrote a word of our AI principles, we asked a room of colleagues two questions: What problems might generative AI solve to improve how government delivers benefits? And, what worries you about using it?

The answers filled a shared board. Caseworkers buried in paperwork need help. Notices are often hard to read. Phone trees bounce people between departments. Records for the same person that don't match across two systems, so someone falls through the gap. A system that sounds confident and is wrong. The same questions answered two different ways on two different days. Personal information processed in new and unfamiliar systems. Over 100 hopes and fears side by side, from more than 20 people across various disciplines in Code for America.

That board became our syllabus.

We could have started by drafting rules and prohibitions. We started here instead, because rules written before that conversation only cover the situations we’d already thought of. We wanted to know what we believed first. 

With so many organizations trying to figure out how to responsibly approach AI while honoring their values and mission, we thought we’d share how we started our journey to help you approach yours.

Here’s how we did it.

With so many organizations trying to figure out how to responsibly approach AI while honoring their values and mission, we thought we’d share how we started our journey to help you approach yours.

We invited anyone who was curious

Our AI working group was open to everyone at Code for America, and we set out to get at least one person from every delivery discipline—areas like quantitative and qualitative research, user experience, engineering, and design. We didn't ask for expertise in AI; we asked for anyone who had tried AI tools themselves and wanted to understand them better.

Staff from 10 different disciplines showed up, representing a wealth of experience from all parts of the product design and delivery landscape. They also brought with them all the experience they had before arriving at Code for America. That means we had people who used to be caseworkers themselves, who have drafted policy in legislatures, and who have seen legacy systems up close while working in government. 

The benefit to bringing that many different perspectives into one room? You get a map of the problem—and many possible solutions—that nobody could have drawn alone.

We created norms to include everyone

Because we didn’t require expertise in AI to join the working group, the entire process was structured for learning. One norm in that setting made a big difference: say so when you don't understand something.

It sounds small, but this norm gave everyone permission to pause the conversation. Some people in that group build machine learning models for a living and others spend their days designing in Figma or annotating policy in Google Docs. Decision rooms can become myopic when everyone’s skillsets look the same. By creating a welcoming environment to ask questions, we ensured that we weren’t leaving people behind, but rather encouraging them to share their perspective. 

AI can be confusing to the vast majority of the population—it’s a highly specialized form of technology that relatively few people know how to build. What will make AI successful in the context of government innovation is an accessibility of understanding. The more people who understand how it works, the more people can shape solutions that use it, and the better those solutions will be for everyone.

From that grounding place of shared involvement, we began learning together.

AI can be confusing to the vast majority of the population—it’s a highly specialized form of technology that relatively few people know how to build. What will make AI successful in the context of government innovation is an accessibility of understanding.

We structured working group sessions to gather information

During our first meeting, we had a brainstorm to build out a set of questions we were interested in exploring, things like:

  • What are some of the problems that generative AI might be able to solve to improve benefits delivery?
  • What are some of the concerns or risks with using generative AI to try and improve government services?
  • What might mitigation look like if models aren’t working as expected?

We met on a recurring basis, with each session having short pre-reading—articles on topics like how AI can demonstrate racial bias based on people’s spoken dialects and how AI chatbots can go wrong in government. We read the latter before a session on accuracy, and discussed as a group what it would take to build and manage a solution to this problem. We then debated whether this was a useful test case for AI, or if the resources used to develop that chatbot might have had a higher rate of return on another project. 

We also brought in six guest speakers, each on a topic the group had asked about. When people wanted to understand how to check a model for accuracy, we found someone who had built evaluations and could show what one actually catches. When the group had questions about privacy, we brought in engineers who had run AI models in secure environments that keep client data from ever reaching a model provider.

Every discussion we had was based in real-world examples, keeping us both up-to-date on the state of AI in government and beyond, and making sure that the principles we were building were directly relevant to the questions our partners were asking. Each session followed the same shape: learn something, build shared vocabulary, and brainstorm against it. 

We debated how, not if

We started from the stated assumption that Code for America will deploy AI solutions to strengthen how we deliver public services—the debate was no longer about if we should use AI, but how to do so responsibly and aligned with our values. Those values are:

  • Listen first
  • Include those who’ve been excluded
  • Act with intention

All of those values were, and are, present in how we think about responsibly implementing AI. 

The debate was no longer about if we should use AI, but how to do so responsibly and aligned with our values.

Throughout the process of working together, we saw measurable growth in working group participants' confidence in identifying opportunities and risks of AI. We polled the working group during the first week, and again at the end of our sessions, and saw improvement on all dimensions we measured. The largest improvement was in participants' self-reported ability to identify the “tactics that Code for America can deploy to mitigate common risks of AI solutions.”

Participants in our AI working group were explicitly expected to carry takeaways back to their discipline and bring their colleagues' perspectives forward—ensuring the group was not a committee, but a distribution mechanism across project teams. Listening, including, and acting formed the core of our work together.

Where we take it from here

The two questions we asked in our first working group session are the ones every team at Code for America is sitting with right now: what problems can we solve in new ways and what are the risks we need to manage?

The work of answering these questions isn't finished—in fact, it never will be. Those questions are the core of what drives Code for America forward, those animating words that support our innovation while keeping the impact on real people front and center.

Technology will keep moving, we'll keep finding new uses for it, and some of what we wrote will need to evolve. We've built a regular review into our planning processes so the principles can change when the work teaches us something.

What we can promise, now and always, is that we'll keep showing our work.

Related stories