TRV-2026-1263Version 1 · Certified
Reason for this version
Certified into the record
Canonical text (the exact bytes fingerprinted)
TRUVACE RECORD VERSION record: TRV-2026-1263 version: 1 kind: certified reason: Certified into the record timestamp: 2026-10-03T06:56:03.596036Z status: published lens: trace sector: science headline: Petal to the metal: The slow road to automating large-scale phenology labeling for herbarium specimens dek: Premise Herbarium specimens represent critical historical records of plant phenology, yet automating annotation of reproductive structures remains challenging given the diversity of floral morphologies, specimen age and quality, and image quality. Methods Here, we present a machine learning pipeline that uses an ensemble modeling approach to detect flowers on herbarium specimens and deliver these data to the phenology research community. After testing multiple strategies for generating training data, we found th… gain_title: An ensemble machine learning pipeline detected flowers on herbarium specimens with relatively strong accuracy and labeled 11.1 million of 22 million records, expanding taxonomic and temporal coverage when integrated into Phenobase. problem_title: The same ensemble pipeline showed moderately high false negative rates for floral structures and only 2.9 million of the 11.1 million flower-labeled records had complete metadata necessary for downstream phenology research. trace_subject: automated detection of flowers on herbarium specimens for phenology research gain_reading: An ensemble machine learning pipeline detected flowers on herbarium specimens with relatively strong accuracy and labeled 11.1 million of 22 million records, expanding taxonomic and temporal coverage when integrated into Phenobase. gain_evidence: modeling pipeline detected present floral structures with relatively strong accuracy | 11.1 million records labeled with flowers present | expands taxonomic and temporal coverage for large-scale phenological analyses problem_reading: The same ensemble pipeline showed moderately high false negative rates for floral structures and only 2.9 million of the 11.1 million flower-labeled records had complete metadata necessary for downstream phenology research. problem_evidence: moderately high false negative rates | only 2.9 million of these contained the complete metadata necessary for downstream phenology research quick_read: By October 2026, researchers described a machine learning pipeline using an ensemble modeling approach to detect flowers on herbarium specimens, addressing challenges from diverse floral morphologies and variable specimen and image quality. After finding expert-curated training data essential, they applied the model to 22 million filtered records and labeled 11.1 million as having flowers present. The work matters because herbarium specimens are critical historical records of plant phenology, and machine-labeled data integrated into Phenobase can expand taxonomic and temporal coverage for large-scale analyses. Uncertainty remains around false negative rates and the fact that only 2.9 million labeled records had complete metadata needed for downstream research. limitation: Performance was constrained by diversity of floral morphologies, specimen age and quality, and image quality, with moderately high false negative rates and dependence on expert-curated training data and complete label digitization. tag: Dual reading key_points: Pipeline used an ensemble modeling approach to detect flowers on herbarium specimens after testing multiple training data strategies. | In-house expert-curated annotations were found essential for producing reliable results. | Final filtered dataset was 22 million records, with 11.1 million labeled as flowers present. | Only 2.9 million labeled records contained complete metadata necessary for downstream phenology research. rundown: Researchers tested multiple strategies for generating training data for flower detection on herbarium specimens and concluded in-house expert-curated annotations were essential for reliable results. Expert validation assessed the ensemble on a filtered final image dataset of 22 million records, resulting in 11.1 million records labeled with flowers present. The authors note only 2.9 million of those labeled records contained complete metadata necessary for downstream phenology research, highlighting the need for full label digitization efforts, and demonstrate integration into Phenobase to expand coverage. sources: - peer_reviewed | Applications in Plant Sciences | https://doi.org/10.1002/aps3.70085 | 2026-10-01 prev: 0000000000000000000000000000000000000000000000000000000000000000
- sha256
- cb5eead4169d63a3de75d40fc557a113cbfe03ae0a8680ff8db60342f43effb4
- previous
- 0000000000000000000000000000000000000000000000000000000000000000
Verify this record
How to verify without trusting this page
Fetch the canonical text of any version from /api/record/TRV-2026-1263 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.
ace