Benchmarking datasets: how Bigpicture helps to build trust in AI
Benchmarking, the systematic evaluation of AI tools against independent reference datasets, is essential for building trust in new technologies. For Bigpicture, benchmarking is not just a technical objective. It's one of the most tangible ways Bigpicture can create lasting scientific and societal impact.
Thanks to the close collaboration between pathologists, AI experts, and industry partners, Bigpicture is ideally positioned to deliver trusted benchmarking datasets for computational pathology.
From hype to proof
AI in pathology is often surrounded by high expectations: faster diagnoses, improved accuracy, and reduced workload. But turning that promise into practice is far from straightforward. Skepticism remains, especially in the medical field. “Soon we’ll have ten algorithms for prostate cancer grading,” says Prof. Jeroen van der Laak, Professor of Computational Pathology at Radboudumc and Project Coordinator of Bigpicture. “Each one will claim to be the best, with excellent numbers on their own datasets. But what does that really tell us? You need an objective reference built on reliable, representative, and independently evaluated data.” That is where benchmarking comes in, transforming claims into evidence, and evidence into trust.
Why benchmarks are harder than you think
It is easy to underestimate what it takes to create a reliable benchmark, but designing one is technically demanding, time-consuming, and costly. “Every benchmark is built for a very specific problem,” Jeroen explains. “Prostate cancer grading requires a different composition of cases than surgical margin assessment. You have to define exactly what the AI is supposed to do, and only then can you decide what data and what ground truth you need.”
But what does ‘ground truth’ actually mean? In pathology, even highly experienced experts may not agree on a diagnosis. Establishing ground truth therefore requires panels, consensus meetings, and sometimes re-scoring cases. It’s complex, but without it, a benchmark cannot be trusted.
To be useful, benchmarks must also reflect clinical reality. “You don’t want everything from one institute,” Jeroen stresses. “The strength of a benchmark comes from including multiple labs, scanners, and protocols. Only then can you trust that a model will work in practice.”
Dr. Jan-Willem Boiten, Partnership Director at Lygature, Co-Lead WP5, adds: “Another major challenge in benchmarking is ensuring strict data protection. Benchmark data must be guarded as securely as the datasets used for training, otherwise, there’s a risk that the same slides are unintentionally used both for algorithm development and for benchmarking. That would undermine the independence and credibility of the benchmark itself.” He continues: “That’s exactly why Bigpicture is uniquely positioned. We bring together hospitals, industry, and regulators across Europe. By combining these perspectives, benchmarks can reflect the diversity you actually see in daily clinical practice.”
Linking benchmarks to Bigpicture’s future
Benchmarking is a crucial topic for Bigpicture and the groundwork is now being laid, as Jan-Willem explains: “If Bigpicture can’t show that the AI we enable actually reaches patients, then what would be the purpose of Bigpicture? We’re not experimenting for the sake of technology. We’re enabling AI that truly makes a difference in clinical practice. Benchmarking is an essential part of making that next step.”
Why benchmarking is key to Bigpicture’s long-term value
Benchmarking is one of the most tangible ways Bigpicture can create lasting scientific and societal impact. “Many similar projects talk about benchmarking as a business model,” says Jeroen. “But they often underestimate the complexity. Bigpicture can make a difference by taking that complexity seriously and building benchmarks together with the community. Among the Bigpicture partners, we have ample experience in organizing grand challenges, which are basically benchmarks disguised as competitions, and even the grandchallenge.org platform is now tightly linked to the Bigpicture repository.”
For industry, benchmarks will provide independent validation to support regulatory approval and market access. For academia, they will offer a way to move research tools closer to clinical translation. For health systems, they will guide procurement and deployment decisions. Jan-Willem underlines the opportunity: “AI models can be built anywhere. But showing, independently, that they are safe and effective in real practice, that’s where Bigpicture can become indispensable.”
How can we as a project ensure the Bigpicture’s infrastructure translates into real impact through benchmarking?
- Defining clinical use cases that should be prioritized (e.g., prostate grading, breast margins).
- Supporting expert involvement to establish consensus ‘ground truth’ for selected tasks.
- Thinking ahead when uploading data: setting aside a portion (for example, 20–25%) for future benchmarking.
Looking ahead
The first benchmarks will not appear overnight. But by protecting data now, defining clear use cases, and involving both pathologists and regulators early, Bigpicture is laying the foundation for benchmarks that will become trusted standards in our field.
As Jeroen concludes: “Without Bigpicture, creating benchmarks would be even more complex. With Bigpicture, we have a unique opportunity to build standards that the entire community can rely on.”
Latest news
20 April, 2026Artificial intelligence is already transforming pathology. But not all pathology questions are the same. In many clinical...
Read more
20 April, 2026When we talk about data upload in Bigpicture, it is often described as a technical process. But...
Read more
20 April, 2026Bigpicture and VICT3R are entering a new phase of collaboration. In a recent discussion, Thomas Steger-Hartmann (Bayer) and...
Read more