The Commonplace
Home Three-study pilot Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Applying ML 'datasheets' to web-archived cultural heritage sharpens provenance and bias transparency but is time- and expertise-intensive, often clashing with archival practice; unless subsidized or adapted, thorough documentation risks concentrating trusted datasets among well-resourced institutions.

Critical questions for documenting collections as data: Provocations from developing datasheets for web archives datasets
Emily Maemura, Helena Byrne · September 10, 2026 · Journal of the Association for Information Science and Technology
openalex descriptive medium evidence 7/10 relevance Summary only summary available; pdf_status=paywall DOI Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

OpenAlex

Latest observation:

  1. Emily Maemura provider ID
  2. Helena Byrne provider ID

Semantic Scholar

Latest observation:

  1. Emily Maemura provider ID
  2. H. Byrne provider ID
Applying ML-style datasheets to UK Web Archive collections improves transparency about provenance and biases but is resource-intensive and reveals conceptual mismatches that require tailored documentation practices and further research.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Abstract Scholarship on documentation of datasets is increasingly concerned with capturing the context, socio‐technical processes, and decisions shaping data. “Datasheets for datasets” is a recent approach developed by researchers in machine learning that foregrounds transparency and uncovering bias through data documentation. Recognizing that machine learning documentation is already being applied for cultural heritage collections as data, we assess the effectiveness of applying datasheets to document web archives datasets. Our study involved community workshops to determine priorities for what to include in a datasheet. We then developed three sample datasheets for datasets from the UK Web Archive, observing and recording the process. Our findings identify the resources needed to create a datasheet, outlining the effort required in terms of time, roles and expertise required. Findings also reveal several notable tensions that emerged in the process, locating mismatches between a framework from machine learning and applications in cultural heritage settings. We develop five provocations to prompt critical consideration of assumptions about how datasets are created and used that might be embedded in any documentation framework. We conclude with an agenda for future research engaging with data documentation in web archives, cultural heritage collections, and data curation.

Summary

Main Finding

Applying machine-learning-style “datasheets for datasets” to web archive collections (UK Web Archive) improves transparency about data provenance and biases but requires substantial time, specialized roles, and adaptation. The ML datasheet framework reveals conceptual mismatches when used in cultural heritage contexts, producing tensions that must be addressed through tailored documentation practices and further research.

Key Points

  • The study tested the ML datasheet approach on cultural-heritage web archive datasets via community workshops and three sample datasheets for UK Web Archive datasets.
  • Workshops identified stakeholder priorities for what documentation should capture; the researchers then created and recorded the process of producing three datasheets.
  • Producing datasheets is resource intensive: it demands significant time, cross-disciplinary expertise (archivists, curators, ML practitioners), and coordination.
  • Several tensions emerged between the assumptions embedded in ML-focused documentation frameworks and realities of cultural heritage collections (e.g., provenance complexity, curatorial practices, access restrictions).
  • The authors formulate five provocations to surface and critique underlying assumptions about dataset creation and use within documentation frameworks.
  • They propose a research agenda to adapt documentation practices for web archives, cultural heritage collections, and data curation more broadly.

Data & Methods

  • Domain: UK Web Archive datasets (cultural heritage/web-archived content).
  • Methods:
    • Community workshops to elicit what stakeholders want included in a datasheet.
    • Development of three exemplar datasheets for real datasets from the archive.
    • Observational recording of the documentation process to measure effort, roles, and challenges.
    • Qualitative analysis to identify mismatches, resource requirements, and themes leading to the five provocations.
  • Outputs: empirical notes on time and personnel required, identification of conceptual tensions, sample datasheets, and a set of provocations and research directions.

Implications for AI Economics

  • Cost of high-quality datasets: Detailed, context-rich documentation increases the labor and time costs of producing datasets, raising the production cost (and potentially market price) of audit-ready or ethically curated datasets.
  • Labor and specialization premium: Documentation requires interdisciplinary labor (archivists, curators, ML engineers), implying higher wages or coordination costs and influencing the structure of dataset production markets.
  • Market differentiation and competitive advantage: Organizations able to invest in thorough documentation may offer higher-trust datasets, creating market segmentation between “documented/trustworthy” and “undocumented” data products.
  • Investment and adoption incentives: The resource intensity may deter smaller institutions or firms from producing documented datasets absent subsidies, standards, or regulation—potentially concentrating high-quality data with well-resourced actors.
  • Regulatory and liability considerations: Robust documentation supports compliance, auditing, and risk management (e.g., bias disclosure), affecting firms’ legal exposure and influencing adoption of best practices as regulatory pressure grows.
  • Externalities and public-good provision: Cultural heritage datasets often have public-value characteristics; documentation costs suggest a role for public funding or collective infrastructure to ensure access to well-documented datasets.
  • Model performance and downstream costs: Better-documented datasets can reduce model development risk (fewer surprises from provenance or bias), lowering downstream mitigation costs, but may also limit dataset reuse if access or licensing constraints are documented.
  • Standardization and coordination effects: Divergence between ML documentation frameworks and cultural-heritage needs indicates potential inefficiencies from one-size-fits-all standards; harmonized, domain-sensitive standards could reduce friction and transaction costs.
  • Trust, adoption, and market demand: Transparent documentation can increase user trust and willingness to pay for datasets, influencing investment in data curation and altering demand for curated vs. scraped/undocumented sources.
  • Research & policy agenda: Economic research is needed on pricing models for documented datasets, subsidies or public provisioning for cultural datasets, incentives for cross-disciplinary labor, and the welfare effects of documentation-driven concentration of high-quality data.

Assessment

Paper Typedescriptive Evidence Strengthmedium — The paper provides primary empirical evidence (community workshops, three exemplar datasheets, and observational records) that supports descriptive claims about feasibility, costs, and conceptual tensions; however the sample is small, qualitative, and non-representative, so findings are suggestive rather than generalizable or causal. Methods Rigormedium — Methods are appropriate for exploratory, practice-focused research (stakeholder workshops, hands-on production of datasheets, observational logging and qualitative analysis). Rigor is limited by small number of exemplars (three datasheets), likely selection/participation biases, and limited detail on analytic coding or triangulation to assess reliability. SampleUK Web Archive domain: stakeholders engaged via community workshops; three real datasets from the UK Web Archive used as exemplars; researchers produced three datasheets and recorded time, roles, and process; qualitative notes and analyses derived from workshop outputs and observational recording. Themesgovernance labor_markets adoption org_design human_ai_collab GeneralizabilitySingle national context (UK) and single domain (web-archived cultural heritage) limit transferability to other countries, data types, or private-sector datasets, Small number of exemplar datasheets (three) limits representativeness across archive collection types and sizes, Workshop participants and stakeholders may not represent the full range of curatorial or ML practitioner perspectives (selection/participation bias), Findings about time/cost are contingent on local institutional practices, resourcing, and access restrictions, Qualitative design means findings are interpretive and may not quantify population-level effects

Claims (6)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Applying ML-style datasheets to UK Web Archive collections improves transparency about dataset provenance and biases. Governance And Regulation positive Transparency of dataset provenance and biases
Reading fidelity high
Study strength medium
n=3
0.18
Producing datasheets for UK Web Archive datasets is resource intensive, requiring substantial time, interdisciplinary expertise, and coordination among archivists, curators, and machine-learning practitioners. Organizational Efficiency negative Labor and coordination requirements for dataset documentation
Reading fidelity high
Study strength medium
n=3
0.18
Stakeholder workshops identified priorities for the information that documentation of cultural-heritage web archive datasets should capture. Governance And Regulation positive Stakeholder-defined documentation requirements
Reading fidelity high
Study strength medium
not reported
0.18
ML-focused dataset documentation frameworks contain assumptions that do not fully align with cultural-heritage web archive collections, particularly regarding provenance complexity, curatorial practices, and access restrictions. Governance And Regulation mixed Fit of a general-purpose ML documentation framework to cultural-heritage datasets
Reading fidelity high
Study strength medium
n=3
0.18
The study formulates five provocations to expose and critically examine underlying assumptions about dataset creation and use in documentation frameworks. Governance And Regulation positive Critical examination of assumptions in dataset documentation
Reading fidelity high
Study strength medium
n=3
five provocations
0.18
The study proposes adapting documentation practices for web archives and cultural-heritage collections rather than relying unchanged on generic ML datasheet frameworks. Governance And Regulation positive Suitability and adaptability of dataset documentation practices
Reading fidelity high
Study strength low
n=3
0.09

Notes