Socio-technical Systems · 2026-03-10

Gene name errors: Lessons not learned

PLOS Computational BiologyOriginal paperMarkdown source
evidencereproducibilitypublic-interest technologyevaluationsdeveloper toolingpublic administration
Key Insight

A decade of documented warnings and nomenclature reforms have not reduced the rate of spreadsheet-induced gene name corruption in published genomics research, demonstrating that knowledge dissemination alone cannot change entrenched data practices; only structural interventions at the software, journal, and training levels can.

Review

Abeysooriya et al. conduct a large-scale longitudinal audit of supplementary Excel files published in PubMed Central between 2014 and 2020, asking whether the well-publicised 2016 report on spreadsheet-induced gene name corruption had any lasting effect on researcher behaviour. The answer is unambiguous: it did not.

Screening 166,139 genomics articles and manually verifying 5,136 suspect files, the authors confirm gene name errors in 30.9% of publications containing supplementary Excel gene lists (3,436 of 11,117), substantially above the 19.6% reported in 2016. The proportion of affected articles remained essentially flat across all seven years, ruling out any secular improvement. The elevated headline figure reflects both a wider sample (all of PMC rather than 18 selected journals) and an improved scanner that now detects five-digit internal date serial numbers, which account for roughly 15.7% of confirmed errors in the sampled subset.

The paper also identifies locale-dependent error modes (gene symbols misread as dates because of Italian, Spanish, Dutch, or Finnish month names), extending the known taxonomy of how spreadsheet internationalisation interacts with biological nomenclature. A counterintuitive finding is a statistically significant positive correlation between journal impact factor and error rate: Cell, Nature, PNAS, and EBioMedicine all exceed 40% affected. The most plausible explanation, unverified in this paper, is that high-impact journals require authors to deposit raw source data, inflating the surface area for errors rather than indicating lower scientific quality.

The central governance-relevant finding is that neither HGNC gene symbol reforms, nor open-source remediation tools (Truke, EscapeExcel, HGNChelper), nor prior publication of the problem have bent the error-rate curve. This generalises beyond genomics: it is a case study in the limits of awareness-based interventions for changing embedded data-handling practices. The paper's prescription (migrate to Python/R notebooks, use LibreOffice when spreadsheets are unavoidable, share genomic data as flat text) is technically sound but underestimates the structural barriers of training, legacy workflows, and journal submission conventions that make behaviour change slow even when intent is present.

For public-sector and research-infrastructure contexts, the implication is that mandating open data deposition without simultaneously mandating format standards and tooling support creates a reproducibility illusion: data are visible but silently corrupted. Journals, funders, and infrastructure operators bear co-responsibility alongside individual researchers.

Key Insight

A decade of documented warnings and nomenclature reforms have not reduced the rate of spreadsheet-induced gene name corruption in published genomics research, demonstrating that knowledge dissemination alone cannot change entrenched data practices; only structural interventions at the software, journal, and training levels can.

Appears in these collections

Continue exploring

Related reviews

More in Socio-technical Systems
Socio-technical Systems · 2026-05-06

Building a Human Resilience Infrastructure for the AI Age

Imagining the Digital Future Center, Elon University

The report's decisive analytical move is to redefine resilience as an institutional property rather than an individual coping skill. Its main governance weakness is that it names contestability, authenticity, literacy, and institutional redesign as necessities without converting them into enforceable decision rights, evidence duties, escalation paths, and redress mechanisms.

Socio-technical Systems · 2026-04-07

AI Assistance Reduces Persistence and Hurts Independent Performance

arXiv preprint

The paper shows that AI assistance is not only a performance aid but a behavioral control surface that can recondition users away from persistence and independent competence. Its governance significance lies in shifting AI evaluation from immediate helpfulness toward measurable autonomy preservation, capability retention, and refusal-to-solve design obligations.

Socio-technical Systems · 2026-03-30

Participatory Unblocking of Blockchain Use Cases: Lessons Learned from the Argentina Onchain Residency

SSRN / BlockchainGov report

The paper’s central value is not that it proves blockchain adoption, but that it reframes blockchain failure as an institutional design problem. Its participatory methodology improves problem selection and contextual fit, but it still stops short of specifying the operational governance needed for legitimate deployment.