Generative AI can draft, summarize, search, translate, classify, and assist with code or customer interactions. Whether that creates value depends on the task, the worker, the workflow, and the controls around the system. A useful enterprise program measures quality and risk alongside speed, and distinguishes assistance from fully automated decisions.

What the evidence supports

JPMorgan Chase: broad access, not a published universal ROI

JPMorgan Chase reported that it launched LLM Suite to more than 200,000 colleagues in 2024. The company described it as access to generative-AI capabilities in a controlled environment designed to protect customer and company data, and as a shared platform on which business teams could develop use cases.

That is evidence of deployment scale and platform strategy. It is not evidence for the previous edition’s claims of a 20% reduction in repetitive work, a 30% reduction in customer wait time, $2 million in data-entry savings, or the attributed executive quote. Those claims have been removed.

Customer support: gains varied by experience

An NBER study examined the staggered introduction of a generative-AI assistant for 5,179 customer-support agents. Access increased issues resolved per hour by 14% on average. The gains were much larger for novice and lower-skilled workers and minimal for the most experienced workers. The study concerns one tool, company, workflow, and period; it should inform a pilot design, not become a default forecast for every department.

Knowledge work: time savings did not automatically reorganize work

A field experiment across 66 firms and 7,137 knowledge workers found that, during the latter half of a six-month experiment, the 80% of treated workers who used the tool spent about two fewer hours on email each week and worked less outside regular hours. The researchers did not detect broader changes in the quantity or composition of tasks from individual access alone. Tool access, adoption, time saved, output quality, and organization-level value are separate measures.

Start with a task, not a transformation slogan

Choose bounded tasks whose inputs, outputs, users, and failure costs are understood. Examples include drafting a first version from approved material, summarizing a controlled document set, retrieving internal policy with citations, suggesting code with tests, or assisting an agent while the employee remains responsible for the interaction.

For each use case, document:

  • the user and decision being supported;
  • allowed data and prohibited data;
  • the model, retrieval sources, tools, and vendors involved;
  • expected benefit and baseline process;
  • foreseeable failures and affected people;
  • required human review and authority to override;
  • quality, security, privacy, fairness, cost, and worker-outcome measures; and
  • conditions for rollback or retirement.

Measure value without hiding trade-offs

DimensionExample measuresCommon mistake
ProductivityCycle time, throughput, time on taskCounting generated words or accepted suggestions as value
QualityTask-specific accuracy, defect rate, rework, groundednessMeasuring speed without checking downstream corrections
RiskPrivacy incidents, policy violations, security findings, harmful errorsTreating a model’s confident answer as verified
PeopleAdoption, workload, autonomy, accessibility, skill developmentAssuming exposure means either full automation or universal benefit
EconomicsTotal cost per successful task, support and review costsIgnoring integration, evaluation, monitoring, and human-review costs

Use a baseline and a comparison group where feasible. Segment results by role, experience, task difficulty, language, and relevant affected groups. Measure long enough to detect novelty effects, workarounds, rework, and changes in model or vendor behavior.

Controls for enterprise use

  • Data governance: classify data, minimize collection, define retention, restrict sensitive inputs, map vendor data flows, and enforce access and deletion requirements.
  • Security: test prompt injection and data exfiltration paths, constrain tool permissions, isolate untrusted content, protect secrets, log actions, and maintain incident response.
  • Output verification: require sources or system records for factual work, test code, and use qualified review for consequential decisions.
  • Employment safeguards: assess discrimination and accessibility risks, provide accommodations, explain monitoring and evaluation practices, and provide meaningful review or appeal where people are affected.
  • Change management: train people on both useful workflows and failure modes; involve workers in workflow design rather than treating resistance as the primary problem.
  • Lifecycle governance: inventory systems and owners, version prompts and models, monitor behavior and cost, reassess material changes, and preserve rollback paths.

How to run a credible pilot

  1. Select a narrow task with enough volume to measure and a manageable failure cost.
  2. Capture the current process, quality, time, cost, and worker experience before introducing AI.
  3. Define acceptance thresholds and prohibited outcomes in advance.
  4. Test with representative users and data, including adversarial and edge cases.
  5. Roll out in stages with human review, telemetry, and a stop mechanism.
  6. Compare results by cohort; investigate who benefits, who does not, and why.
  7. Scale only when the combined evidence on quality, risk, economics, and people supports it.

Work is more likely to change by task than vanish by job title

The International Labour Organization’s 2025 exposure analysis found that one in four workers is in an occupation with some generative-AI exposure, while only 3.3% of global employment is in its highest exposure category. Because occupations contain many tasks that still require human input, the report identifies job transformation as the more likely broad effect. Exposure is not the same as adoption, automation, productivity, or job loss, and impacts differ by occupation, country, and gender.

Bottom line

Enterprise generative AI is neither a guaranteed productivity engine nor merely a chatbot rollout. The strongest programs select specific tasks, establish evidence before scaling, protect people and data, and keep accountable humans in control of consequential work. The question is not simply whether employees can access a model, but whether the redesigned system produces better outcomes under real operating conditions.

Primary and authoritative sources