120,000 skin pathology slides: Region Östergötland prepares its Bigpicture contribution
What does it take to turn 120,000 clinical pathology slides into data that researchers can actually use? As Region Östergötland prepares one of Bigpicture’s major skin pathology contributions, the work shows why building a valuable research dataset involves much more than moving images from one system to another.
Bigpicture is operational and already contains data contributed by partners. Region Östergötland is now preparing its own skin pathology contribution, expected to be organized into 16 related datasets. Anna Bodén, Bigpicture Work Package 3 co-lead and node coordinator in the Skin node, says: “It’s not just to upload data. The aim is to prepare material that researchers can understand, select and ultimately use.”
Inside the 120,000-slide collection
The planned contribution covers pathology material from 2019 to 2025, based on 21 diagnosis registrations in Region Östergötland’s laboratory information system. The collection includes malignant melanoma alongside other malignant and benign skin lesions that may be clinically or morphologically relevant when distinguishing melanoma from other conditions. Rather than making the material available as one very large dataset, the plan is to divide it into related collections. This should allow researchers to focus on a specific diagnosis while still identifying other datasets that could be useful for training, comparison or validation.
The contribution will include whole-slide images and the metadata required within Bigpicture. For malignant melanoma, the team is also preparing additional information from pathology reports to provide clinical context at case level. Different stains will be retained, including H&E and immunohistochemistry, giving researchers more flexibility to select material relevant to their own questions. The collection will be anonymized and will not contain handcrafted annotations. Its value instead lies in carefully validated and curated metadata that allows researchers to understand and select relevant material.
Why a melanoma dataset needs more than melanoma
The original motivation came from a practical pathology challenge: could AI help identify and triage true melanoma cases? To investigate that kind of question, an algorithm needs to be able to handle the broader clinical landscape of skin lesions, not just confirmed melanomas. It also needs relevant alternatives: lesions that may resemble melanoma clinically or morphologically. The team therefore worked with dermatopathology expertise to determine which diagnoses should be included. The result is a broader collection that can support comparison between different types of skin pathology rather than a narrowly defined melanoma dataset. This also leaves room for research questions beyond the original use case.
Testing AI across institutions
Potential applications discussed by the team include detection tasks, biomarker prediction and quantification. There is also value in combining the contribution with data from other centers. AI developed on material from one institution may perform differently when applied elsewhere. Access to datasets from multiple organizations can help researchers investigate this variation and work toward more robust models.
Region Östergötland has already been collaborating with the team at Semmelweis to align the diagnoses included in their datasets, helping create opportunities for external validation across institutions. That is one of the advantages of contributing to Bigpicture: individual datasets become part of a larger European resource that can support research across centers.
What researchers will be able to work with
- A broad skin pathology collection rather than one narrowly defined use case;
- Related datasets that can be selected individually or used in combination;
- Whole-slide images with accompanying, curated metadata;
- H&E and immunohistochemistry material;
- Opportunities for AI development, comparison, external validation and multicenter research.
Additional curation or pathology expertise may still be needed depending on the research question.
120,000 slides means 120,000 data challenges
The team first had to understand the source material: how cases were distributed across diagnoses and years, how many slides were associated with each case, and where metadata needed to be reviewed or matched. Cases with unusually high slide numbers, above the 95th percentile, were excluded from the planned extraction.
Experience with the extraction tools helps, but limitations in laboratory information systems, particularly around data granularity and standards, mean that manual input is still necessary. Pathologist involvement in validation and curation at specimen and slide level therefore remains substantial, even where parts of the process can be supported by automation. As Anna explains: “Most of the data needs to be curated, as this is a data-driven project. If the input to Bigpicture is not of good quality, the AI results will be mostly unpredictable. Poor-quality training data leads to poor-quality AI outputs. It’s a classic case of garbage in, garbage out.”
The work has also raised practical questions about how the material should be divided and presented, including the design of the anonymization procedure in line with the ethical approval for submission and research use of anonymous data. Within Bigpicture, work is also ongoing to harmonize and map data to different standards.
Seven lessons for future contributors
- Start with the clinical need. Use it to define the intended research value and consider what future users may need before defining the extraction.
- Understand your information and data sources. Review case volumes, metadata quality, slide numbers and diagnostic categories, and map the extraction pathway to understand what is feasible with the resources available.
- Expect curation at scale. More data also means more validation, metadata matching and organizational work.
- Think about access as well as upload. The most convenient structure for contribution is not necessarily the most useful structure for research.
- Use clinical expertise. Selecting relevant comparison material requires domain knowledge, not only technical extraction criteria.
- Involve data stewards and analysts who understand your internal systems. Extracting clinical data at this scale requires a combination of pathology, data and systems expertise.
- Make sure the necessary approvals are in place. An approved ethical permission and local data management plan are essential.
Researchers don’t have to wait
Region Östergötland’s collection is still being prepared. Alongside the curation and validation work, the team has been adapting extraction tools and pathways to the newly developed data standard. Researchers, however, do not need to wait for every planned contribution to arrive. Bigpicture is already operational and contains data from other partners. New collections will continue to broaden that resource, adding more diagnoses, institutions and material for research.
The opportunity is therefore both to contribute more useful data and to start using the data that is already available. Region Östergötland’s contribution is one part of that growing resource, and its preparation shows why making data available is only the first step. Making it usable is what gives it value.
Latest news
What does it take to turn 120,000 clinical pathology slides into data that researchers can actually use...
Read more
Being unique is one thing. Being valuable enough for organizations to rely on is another. Dr. Page...
Read more
The biggest hurdle for pathology AI may no longer be building the technology, but proving it works...
Read more