Background
The current Apis mellifera reference (Amel_HAv3.1) is built from a single individual and systematically misses structural variation that segregates between subspecies. Conservation genetics, breeding programs, and pesticide-response studies all depend on calls made against this single haplotype — and many of those calls disagree across cohorts in ways that are not biological.
Objective
We are constructing a graph reference for the honeybee from 240 haplotype-resolved long-read assemblies spanning four subspecies (A. m. mellifera, A. m. ligustica, A. m. carnica, A. m. iberiensis). The deliverable is a versioned, downloadable graph (in GFA) plus a projection layer that lets short-read pipelines re-call variants against it without changing their tooling.
Aims
- Assemble haplotype-resolved long-read genomes for 240 individuals from existing public and partner-contributed Nanopore R10 data.
- Build a population graph with
graphmap2+pansift, with explicit subspecies tagging on graph paths. - Quantify the diagnostic gain for variants of known phenotypic effect (Varroa-resistance loci, viral-tolerance QTLs) when calling against the graph versus the linear reference.
Honeybee assemblies are contributed by external partners including a university group, a national agricultural research institute, and an apiculture consortium. The graph and projection layer are released under CC-BY 4.0 with code under MIT.