Stop Starting with Ministries: A New Map for Where Government AI Actually Works

Stop Starting with Ministries: A New Map for Where Government AI Actually Works

A landmark study presented in Madrid argued that the reason 40% of public-sector AI projects fail is structural — and offered a function-based way to fix it. 

Manuel Kilian came to the Main Stage with a number designed to focus attention. Roughly nine in ten citizens and businesses say they would be willing to use AI agents in their dealings with public administration. Roughly nine in ten public institutions say they plan to explore or deploy agentic AI within two to three years. And yet, by his estimate, around 40% of those projects may be discontinued because they fail – a failure rate that does more than waste money. It corrodes the very public trust that makes digital government possible. 

As Managing Director of the Global Government Technology Centre in Berlin, Kilian was in Madrid to present the 2026 Global State of GovTech Report, produced with Capgemini and co-published with the World Economic Forum. Its central theme was agentic AI in government, and its central argument was uncomfortable: the reason so many projects fail is not the technology. It is the shape of the government itself. 

A Structural Mismatch 

Agentic AI, he explained, combines the “brains” of large language models with the “hands” of tools and workflow execution. It operates across end-to-end processes. Governments, by contrast, organize themselves around ministries, departments, org charts, and lines of political accountability. The two do not fit. Drop an AI agent designed to run a workflow into an institution built around silos, and it will struggle, not because the model is weak, but because the structure resists it. 

His proposed corrective was simple to state and demanding to implement: begin not with the institutions that govern, but with the functions government must perform. Instead of asking how a ministry can “adopt AI,” ask which repeatable, high-volume workflows cut across agencies and are common to public administration everywhere. 

Decomposing The Work of Government 

The report’s method, developed over roughly six months, was to identify universal government workflows that are discrete, high-impact, and common across jurisdictions. Kilian used document validation as an illustration: receiving documents, verifying authenticity, checking completeness, and extracting information. Once a service is broken down this way, specific AI agents can be mapped to specific steps, making implementation modular and realistic rather than monolithic and aspirational. Trying to “apply agentic AI” to something as broad as childcare support in one leap is a recipe for failure. Breaking it into reusable components, like document validation, case tracking, benefit calculation, payment processing, queue management, is not. 

The research identified 70 core government functions that together describe most of what governments do, grouped into core service delivery, policy and governance, and enabling infrastructure. One of Kilian’s more striking observations was that the biggest opportunities may not lie in visible, citizen-facing interactions at all, but in the unglamorous back office. He singled out tender preparation and awarding in public procurement, a process of enormous economic weight given that procurement accounts for a substantial share of global GDP. Even modest automation gains there could have outsized impact. 

A Topography of Potential 

To turn this into a decision tool, the report assesses each of the 70 functions on two axes.  

The first is agentic AI potential: can the function be automated, should it be handled by an agent rather than simpler automation, and is the workflow high-volume and high-impact?  

The second is implementation complexity: data requirements and quality, technical integration burden, cross-agency dependencies, ethical considerations, and regulatory constraints.  

Plotting functions across these axes produces what Kilian called a “topography of potential,” dividing opportunities into high-, medium-, and low-readiness zones. 

The high-readiness zone, where strong AI potential meets relatively low complexity, is where governments should begin. The framework’s value, he argued, is discipline. Rather than launching scattered pilots, administrations can identify the five to ten workflows most worth pursuing, check whether their current initiatives are aimed at the right targets, and publicly signal priorities to the market so that vendors and delivery partners can align with genuine demand. He was careful to add that countries should adapt the scoring to their own realities, since complexity factors such as data quality vary enormously by jurisdiction. 

Key Takeaways 

  1. Most public-sector AI projects fail for structural, not technical, reasons. Agentic AI runs on workflows. Governments are organized around silos. The mismatch is the problem. 
  1. Start with functions, not institutions. The report identifies 70 core government functions common across jurisdictions – a more useful starting point than any single ministry’s org chart. 
  1. Decompose services into reusable components. Breaking a service into document validation, case tracking, benefit calculation, and the like makes AI implementation modular and achievable. 
  1. The biggest wins may be in the back office. Unglamorous, high-volume processes like procurement and tendering offer outsized economic returns from even modest automation. 
  1. Prioritize by readiness. Assess each function on AI potential and implementation complexity, and start in the high-readiness zone where strong potential meets manageable complexity. 

The GovTech 4 Impact World Congress returns in 2027. To stay connected and be the first to hear about what comes next, visit g4i-congress.com and follow us on social media.  #G4I2027