Data upload in action: Helsinki University Hospital (HUS) leads by example
In the Bigpicture project, the uploading of pathology data is more than a milestone. It's a signal that Europe’s vision for a shared research infrastructure is taking shape. One of the first partners to successfully complete a real data upload is the team from the Helsinki Biobank (HUS). Their work not only demonstrates what’s possible, but it also provides valuable lessons for others preparing to upload.
We spoke to Prof. Olli Carpen (WP3, Lead Honest Broker Task Force and Lung Node), Yossra Zidi (WP3, Lung Node Coordinator and Pathologist), and Kimmo Ala-Kulju (Data Scientist) about their role in Bigpicture, the journey toward their first successful upload, and what they’ve learned along the way.
A strong foundation in digital pathology
HUS has long been a leader in health data infrastructure. With over 20 years of experience in digitalized pathology database and a biobank that has been investing in data integration since its inception, Helsinki was well-positioned to contribute to Bigpicture. “This was a perfect project for us to advance digital pathology,” says Olli. “We’ve built a strong datalake combining pathology, lab values, treatment records, and more. That made it possible to link samples with rich clinical metadata, a key requirement for Bigpicture.”
Their role spans Work Package 3, where they manage the lung node, and are also helping to shape the honest broker mechanism, with active member Minttu Sauramo contributing to its practical implementation. This dual contribution, both technical and governance-related, puts them at the heart of the project’s infrastructure.
From planning to uploading: a team effort
While the digital foundation was strong, uploading a dataset proved to be a complex process, requiring collaboration across pathology, data science, biobank coordination, histology, and legal teams. “It’s been a long road,” says Yossra, who has been involved in Bigpicture since 2021 and played a key role in designing the metadata and sample selection process. “There was nothing at the start, no structure, no templates. We helped build them from scratch through the metadata task force.”
Their workflow was carefully aligned with the biobank’s normal operations to avoid overburdening staff. From defining which cases to include, to physically retrieving and scanning over 8,000 slides, every step was documented and iterated. Key to their success was designing a process that others in the biobank could follow without stepping too far outside their usual duties.
“That’s been a big factor,” Yossra notes. “We tried to keep it as close as possible to our usual workflow. That helped get internal buy-in and made it easier to repeat.”
Technical complexity: mapping, metadata, and automation
Once slides were selected and scanned, the task shifted to metadata transformation. Kimmo developed automated workflows to extract structured data from the hospital’s datalake and map it to the required XML format. “We use automation where we can,” he explains, “but there’s still some manual mapping, especially translating diagnosis terms from Finnish and matching them to SNOMED CT codes.”
Over time, the team built a growing “mapping database”, allowing previously mapped items to be reused and making each new dataset faster to prepare. “The first dataset was a big one,” Kimmo says. “It took time, but now that the scripts are ready and the mapping is accumulating, the next ones are already going much faster.”
Lessons learned, and shared
As one of the first to successfully complete an upload, the team from HUS is keen to support others. “My recommendation to other data submitters is to ensure that the team includes at least a pathologist and a data scientist working closely together at every step,” says Yossra. Other key advice from the HUS team:
- Start small: “Don’t upload everything in one go. Start with a manageable dataset.”
- Use structured data: “Structured tables are much easier to map and verify than free text.”
- Align with your workflow: “Involve the right people early and fit into existing routines.”
- Automate where you can: “Even partial automation can significantly reduce workload over time.”
- Test with mock datasets: “This helped us prepare without needing full authorizations upfront.”
Looking ahead: what are the next steps for HUS?
The Helsinki team is now working on their second and third datasets, and collaboration continues with CSC (ELIXIR node) and Work Package 2 to streamline future uploads. Weekly “virtual coffees” have helped maintain close contact, and the team is already thinking beyond upload and more towards visibility and impact. “Seeing the landing page, visualizing the dataset itself, and how our dataset is used, that’s what will be most motivating,” says Yossra. “Right now, it’s still abstract. We need to visualize what we’ve achieved.”
A realistic but inspiring example
Uploading a dataset is no small task. It requires technical skill, coordination, patience, and teamwork. The HUS team shows it can be done, and that it gets easier with each step.
Latest news
20 April, 2026Artificial intelligence is already transforming pathology. But not all pathology questions are the same. In many clinical...
Read more
20 April, 2026When we talk about data upload in Bigpicture, it is often described as a technical process. But...
Read more
20 April, 2026Bigpicture and VICT3R are entering a new phase of collaboration. In a recent discussion, Thomas Steger-Hartmann (Bayer) and...
Read more