0 cumulative citations
View corpus contextApplying ML 'datasheets' to web-archived cultural heritage sharpens provenance and bias transparency but is time- and expertise-intensive, often clashing with archival practice; unless subsidized or adapted, thorough documentation risks concentrating trusted datasets among well-resourced institutions.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
0 cumulative citations
View corpus contextAbstract Scholarship on documentation of datasets is increasingly concerned with capturing the context, socio‐technical processes, and decisions shaping data. “Datasheets for datasets” is a recent approach developed by researchers in machine learning that foregrounds transparency and uncovering bias through data documentation. Recognizing that machine learning documentation is already being applied for cultural heritage collections as data, we assess the effectiveness of applying datasheets to document web archives datasets. Our study involved community workshops to determine priorities for what to include in a datasheet. We then developed three sample datasheets for datasets from the UK Web Archive, observing and recording the process. Our findings identify the resources needed to create a datasheet, outlining the effort required in terms of time, roles and expertise required. Findings also reveal several notable tensions that emerged in the process, locating mismatches between a framework from machine learning and applications in cultural heritage settings. We develop five provocations to prompt critical consideration of assumptions about how datasets are created and used that might be embedded in any documentation framework. We conclude with an agenda for future research engaging with data documentation in web archives, cultural heritage collections, and data curation.
Summary
Main Finding
Applying machine-learning-style “datasheets for datasets” to web archive collections (UK Web Archive) improves transparency about data provenance and biases but requires substantial time, specialized roles, and adaptation. The ML datasheet framework reveals conceptual mismatches when used in cultural heritage contexts, producing tensions that must be addressed through tailored documentation practices and further research.
Key Points
- The study tested the ML datasheet approach on cultural-heritage web archive datasets via community workshops and three sample datasheets for UK Web Archive datasets.
- Workshops identified stakeholder priorities for what documentation should capture; the researchers then created and recorded the process of producing three datasheets.
- Producing datasheets is resource intensive: it demands significant time, cross-disciplinary expertise (archivists, curators, ML practitioners), and coordination.
- Several tensions emerged between the assumptions embedded in ML-focused documentation frameworks and realities of cultural heritage collections (e.g., provenance complexity, curatorial practices, access restrictions).
- The authors formulate five provocations to surface and critique underlying assumptions about dataset creation and use within documentation frameworks.
- They propose a research agenda to adapt documentation practices for web archives, cultural heritage collections, and data curation more broadly.
Data & Methods
- Domain: UK Web Archive datasets (cultural heritage/web-archived content).
- Methods:
- Community workshops to elicit what stakeholders want included in a datasheet.
- Development of three exemplar datasheets for real datasets from the archive.
- Observational recording of the documentation process to measure effort, roles, and challenges.
- Qualitative analysis to identify mismatches, resource requirements, and themes leading to the five provocations.
- Outputs: empirical notes on time and personnel required, identification of conceptual tensions, sample datasheets, and a set of provocations and research directions.
Implications for AI Economics
- Cost of high-quality datasets: Detailed, context-rich documentation increases the labor and time costs of producing datasets, raising the production cost (and potentially market price) of audit-ready or ethically curated datasets.
- Labor and specialization premium: Documentation requires interdisciplinary labor (archivists, curators, ML engineers), implying higher wages or coordination costs and influencing the structure of dataset production markets.
- Market differentiation and competitive advantage: Organizations able to invest in thorough documentation may offer higher-trust datasets, creating market segmentation between “documented/trustworthy” and “undocumented” data products.
- Investment and adoption incentives: The resource intensity may deter smaller institutions or firms from producing documented datasets absent subsidies, standards, or regulation—potentially concentrating high-quality data with well-resourced actors.
- Regulatory and liability considerations: Robust documentation supports compliance, auditing, and risk management (e.g., bias disclosure), affecting firms’ legal exposure and influencing adoption of best practices as regulatory pressure grows.
- Externalities and public-good provision: Cultural heritage datasets often have public-value characteristics; documentation costs suggest a role for public funding or collective infrastructure to ensure access to well-documented datasets.
- Model performance and downstream costs: Better-documented datasets can reduce model development risk (fewer surprises from provenance or bias), lowering downstream mitigation costs, but may also limit dataset reuse if access or licensing constraints are documented.
- Standardization and coordination effects: Divergence between ML documentation frameworks and cultural-heritage needs indicates potential inefficiencies from one-size-fits-all standards; harmonized, domain-sensitive standards could reduce friction and transaction costs.
- Trust, adoption, and market demand: Transparent documentation can increase user trust and willingness to pay for datasets, influencing investment in data curation and altering demand for curated vs. scraped/undocumented sources.
- Research & policy agenda: Economic research is needed on pricing models for documented datasets, subsidies or public provisioning for cultural datasets, incentives for cross-disciplinary labor, and the welfare effects of documentation-driven concentration of high-quality data.
Assessment
Claims (6)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Applying ML-style datasheets to UK Web Archive collections improves transparency about dataset provenance and biases. Governance And Regulation | positive | Transparency of dataset provenance and biases |
Reading fidelity
high
Study strength
medium
|
n=3
|
| Producing datasheets for UK Web Archive datasets is resource intensive, requiring substantial time, interdisciplinary expertise, and coordination among archivists, curators, and machine-learning practitioners. Organizational Efficiency | negative | Labor and coordination requirements for dataset documentation |
Reading fidelity
high
Study strength
medium
|
n=3
|
| Stakeholder workshops identified priorities for the information that documentation of cultural-heritage web archive datasets should capture. Governance And Regulation | positive | Stakeholder-defined documentation requirements |
Reading fidelity
high
Study strength
medium
|
not reported
|
| ML-focused dataset documentation frameworks contain assumptions that do not fully align with cultural-heritage web archive collections, particularly regarding provenance complexity, curatorial practices, and access restrictions. Governance And Regulation | mixed | Fit of a general-purpose ML documentation framework to cultural-heritage datasets |
Reading fidelity
high
Study strength
medium
|
n=3
|
| The study formulates five provocations to expose and critically examine underlying assumptions about dataset creation and use in documentation frameworks. Governance And Regulation | positive | Critical examination of assumptions in dataset documentation |
Reading fidelity
high
Study strength
medium
|
n=3
five provocations
|
| The study proposes adapting documentation practices for web archives and cultural-heritage collections rather than relying unchanged on generic ML datasheet frameworks. Governance And Regulation | positive | Suitability and adaptability of dataset documentation practices |
Reading fidelity
high
Study strength
low
|
n=3
|