Arthur Goldstuck | CEO | World Wide Worx | Editor-in-Chief | Gadget.co.za | mail me |
Artificial intelligence (AI) systems now pass bar exams. They write code. They even mimic empathy. Yet few answer a simpler question: are they helping people live better lives? Put differently, what good is AI if it does not improve human wellbeing?
That question now has a practical test. The Flourishing AI Benchmark (FAI) offers a new way to evaluate systems. Faith-based technology company Gloo developed the benchmark. Instead of measuring how clever a model appears, it measures how much good it can do. In other words, it directly confronts the question of what good is AI?.
Measuring wellbeing, not cleverness
The FAI avoids the standard approach used in most AI testing. It does not prioritise speed, logic, or technical optimisation. Instead, it evaluates responses across seven dimensions of human wellbeing. These include character, relationships, happiness, meaning, health, finances and faith. The benchmark reframes performance around impact. It asks, again, what good is AI if it ignores these dimensions?
The concept of human flourishing is not new. The idea stretches back centuries. Aristotle explored it first. Modern science continues to study it today. The benchmark also draws on frameworks like the Global Flourishing Study. That study includes data from more than 200,000 people across over 20 countries.
– Pat Gelsinger, Executive Chairman and Head of Technology at Gloo
Gelsinger brings deep technical credibility to this effort. He previously served as CEO of Intel and VMware. He also helped architect foundational technologies such as USB and Wi-Fi standardisation. Now, he proposes a new standard for the AI age. This standard focuses less on capability and more on consequence.
From technical performance to human impact
Traditional metrics still matter. However, they focus only on technical capability. They measure response speed or correctness in defined areas like mathematics. That focus is too narrow. The industry must expand the conversation. It must ask whether AI responses actively promote human wellbeing. That shift defines the real debate about what good is AI.
Avoiding harm is not enough. Systems must demonstrate a positive contribution. It isn’t just the absence of bad. We must show the presence of good. That distinction underpins the entire benchmark.
To test this idea, researchers evaluated more than two dozen advanced large language models. They answered 1,229 questions aligned to the seven wellbeing dimensions. The design ensured both breadth and depth.
How the benchmark works
The questions included objective multiple-choice items. They also included subjective, scenario-based prompts. Other large language models evaluated the responses. Each acted as a judge. Each judge adopted a specific expert persona linked to the dimension under review.
These judges followed a structured rubric. The rubric included 25 scoring criteria. When a response touched multiple dimensions, additional judges assessed it. This method ensured consistency. It also reinforced the benchmark’s holistic intent.
The FAI represents more than a scoring tool. It challenges the industry directly. Our goal is to make AI better across the board. The benchmarks exist to encourage progress and measure it. At present, no model meets the full threshold for human flourishing.
A challenge to the AI industry
Some models perform well in areas like finance or health. However, many struggle with meaning or faith. That imbalance matters. If it’s not supporting human flourishing, the engineering isn’t done. Such failures are bugs that require fixing.
Comparisons are drawn to other technologies. If a self-driving car crashes too often, engineers return to work. If a humanoid robot endangers people, designers revise it. Likewise, if AI systems fail to reflect human values, developers must improve them. Otherwise, the question of what good is AI remains unanswered.
Accountability is emphasised. The FAI helps create shared responsibility and clearer ethical standards. These standards must remain open, adaptable and scalable. Still, ethics cannot sit with technologists alone.
Beyond technologists and toward shared responsibility
Ethical AI requires diverse voices. These voices must include ethicists and psychologists. They must also include faith leaders across traditions. Everyday users matter as well. AI remains a tool. It should never replace humans.
Every AI response should make that limitation clear. Systems should guide people toward better outcomes. That includes better relationships with other humans. Here, the benchmark returns to its core concern: what good is AI if it weakens human connection?
With proper safeguards, AI can support care. It should not supplant it. There is a need for context, privacy, guardrails, security and access. Without these, assistance becomes risk.
An ambitious but practical goal
AI adoption remains early. Despite this, society already debates its impact on wellbeing. This is progress, and this is a major win.
The long-term goal is ambitious. Yet it remains practical. We can teach and educate every child on the planet. By doing so, AI could help lift people out of extreme poverty. Eventually, it could help end it. That vision offers a clear answer to what good AI is.


























