How Goldman Sachs built an AI usage policy that employees actually followed
Most corporate AI policies sit in a shared drive and change nothing. Goldman Sachs took a different path, and the mechanics of how they did it offer a transferable model for any team serious about governing AI in practice.
Neo NeumannAI Practice LeadSeptember 1, 2026Listen to the podcast
4 min
By mid-2025, Goldman Sachs had deployed its internal AI assistant, GS AI, to most of its roughly 45,000 employees globally. The tool, built on a combination of proprietary infrastructure and models from providers including OpenAI, gave analysts, developers, and client-facing staff access to generative AI for drafting, summarizing, and coding tasks. The problem Goldman faced was not access. It was behavior. Employees were using AI inconsistently: some sharing client data in prompts, others using personal ChatGPT accounts on work problems because the approved tool felt slower, and many simply unaware of where the lines were.
The compliance and technology teams recognized they had a governance gap. A policy document, however well-written, would not close it. What they needed was a set of mechanisms that made the right behavior the path of least resistance.
What they did
Goldman's approach had four concrete components, each targeting a different failure mode.
First, they built the policy into the tool interface, not a separate PDF. When a user starts a prompt in GS AI, a brief contextual reminder appears if the system detects patterns associated with sensitive data, such as client names or account numbers in the input field. This is not a hard block; it is a friction point that prompts the user to reconsider. The distinction matters. Hard blocks create workarounds. Friction creates pauses.
Second, they ran role-specific training rather than a single all-hands briefing. A junior analyst in equity research received a 45-minute module focused on three scenarios: summarizing earnings calls, drafting client memos, and handling confidential financial projections. A software engineer received a different module centered on code generation, IP considerations, and the risk of leaking proprietary algorithms into public model APIs. The training was short, scenario-based, and tied directly to work the employee already did. Generic AI ethics content was kept minimal.
Third, Goldman used behavioral data from the platform to identify where the policy was failing. If a significant number of engineers were consistently bypassing GS AI for a particular task and using external tools instead, that was treated as a signal that the internal tool had a gap, not that the employees were rogue. This feedback loop led to several product iterations within the first year of broad deployment, including improvements to code completion speed, which had been a primary driver of external tool use.
Fourth, managers were given a one-page "AI conversation guide" designed to help them talk about AI use with their teams in regular check-ins, not in formal review settings. The goal was to normalize the conversation about which AI behaviors were acceptable and which were not, so that questions did not accumulate into undisclosed habits.
The results
Goldman has not published a detailed external breakdown of policy compliance metrics, so specific figures should be treated as approximate where noted. Internal reporting cited by Bloomberg in late 2025 indicated that usage of GS AI had grown to the point where the firm was processing millions of AI-assisted interactions per week, with a measurable reduction in employees using unapproved external tools for work tasks. The compliance team reported a reduction in data handling incidents related to AI use, though the baseline and exact percentage improvement have not been disclosed publicly.
What is documented: Goldman's developers attributed a meaningful share of their productivity improvement in software engineering to the AI coding assistant, and the firm's technology leadership stated publicly at their 2025 investor day that AI tooling was contributing to engineering efficiency in ways they expected to quantify more precisely over the following 12 months.
The behavioral shift that is probably most significant, and hardest to measure, is the normalization of the governance conversation itself. Employees at Goldman began treating AI tool selection and prompt behavior as a professional skill with real stakes, similar to how they already treated data classification or communication compliance. That shift does not happen from a policy document. It happens from consistent reinforcement through tooling, training, and management conversation.
What transfers
The Goldman model contains four lessons that apply across organizations at very different scales.
The first is that policy enforcement and product design are the same problem. If your approved AI tool is noticeably worse than the free consumer alternative for the task your employees actually need to do, they will use the free tool. Closing that gap is a governance act, not just a product decision.
The second is role specificity. A policy that tries to cover every employee in every function produces training content that feels irrelevant to everyone. SegmentingSegmentingDividing a market into distinct groups of customers who share similar needs, characteristics or behaviours, so each group can be served with a tailored approach.View full definition → by job function and writing scenarios in the employee's own workflow language is more work upfront and produces better compliance. Two or three genuinely relevant scenarios beat twenty generic principles.
The third is using behavioral data as a governance signal. Most organizations have access to usage logs from their AI platforms and treat them purely as a billing or capacity metric. The more useful read is: where are employees not using the approved tool, and why? That question leads to better policy, better tooling, and more honest conversations about what the policy was actually asking employees to do.
The fourth is manager activation. Compliance teams cannot watch every interaction. Managers can normalize the right conversations in their existing rhythms. One page of concrete talking points, tied to scenarios the manager's team actually encounters, is more likely to be used than a detailed governance framework.
The place where most organizations differ from Goldman is resources. Goldman has a large compliance function, significant technology investment, and a regulatory environment that made AI governance a genuine priority, not a nice-to-have. Smaller teams working in less regulated industries will need to make harder choices about which of these levers to pull first. In that case, role-specific scenario training is the one to start with, because it costs the least and changes mental models the most.
A policy that changes behavior is a design problem, not a writing problem. Goldman figured that out early and built the infrastructure to match. The specifics of their tooling will not translate directly, but the logic of making compliant behavior the convenient choice will apply to any team deploying AI at scale.
Go deeper
The lessons that take this article further, free to read.
- 1Privacy and confidential data: what not to pasteResponsible & trustworthy AI
- 2Ethics and responsible use at workResponsible & trustworthy AI
- 3Governance and the EU AI Act: the toplineResponsible & trustworthy AI
- 4Sharing, governance, and custom gpts at workChatGPT & the OpenAI ecosystem
- 5Verifying outputs: trust but checkAI in daily work
Finished reading?
Validate your read to earn XP and feed your radar.