Google Research study finds transfer learning from European GWAS data improves genomic risk prediction in small Japanese cohorts but degrades accuracy once target cohort size grows
Google Research published a study evaluating polygenic risk score (PRS) transferability across eight clinical traits using UK Biobank (European) and Biobank Japan cohorts, comparing three methods (UKB discovery GWAS + elastic net, meta-analysis + elastic net, PRS-CSx). Findings: pooling European GWAS data with target-population data boosts PRS accuracy when target sample sizes are small, but beyond a crossover point (around 15,000 BBJ samples for most traits, higher for traits with conserved genetic architecture across populations), adding more European data degrades predictive performance compared to training on target-population data alone.
Entities: Google Research, Joey Poomarin Phloyphisut, Cory McLean, UK Biobank, Biobank Japan, PRS-CSx
1 primary
What happened
Google Research published a study testing whether adding European genetic data (UK Biobank) improves polygenic risk score accuracy for a Japanese cohort (Biobank Japan) across eight clinical traits, using three methods including PRS-CSx. The result: pooling in European data helps when the Japanese sample is small, but past a crossover point, roughly 15,000 Biobank Japan samples for most traits (higher for traits with more conserved genetics across populations), adding more European data actually makes predictions worse than just using target-population data alone.
Why it matters
This gives researchers building genomic risk models for non-European populations a concrete, quantified rule of thumb for when to stop leaning on European datasets and switch to local data, which could improve prediction accuracy in real projects. The impact is confined to methodology for researchers and companies building PRS tools; polygenic risk scores are still not widely used in clinics, so nothing here changes patient care today.
What is noise
The framing suggests this addresses low clinical adoption of PRS in underrepresented populations, but the study is a methods ablation, not a clinical validation or deployment, and adoption barriers (regulatory, cost, clinical utility) are untouched by this work. The core problem, European-biased GWAS data, and the general idea of transfer learning to correct for it, are already well known; the actual new contribution is narrower: a specific sample-size threshold and a counterintuitive crossover effect.
Watch next
- 01Whether the paper is published in or submitted to a peer-reviewed journal, and how reviewers treat the ~15,000-sample crossover claim.
- 02Whether other groups replicate the crossover finding in different non-European cohorts (e.g. African, South Asian ancestry biobanks), which would test how generalisable the threshold really is.
- 03Whether PRS-CSx or similar tools update their default guidance or documentation to reflect this sample-size-dependent recommendation, indicating practical uptake beyond the original paper.
Coverage
1 storyMore capability signals
Full feed →- AI systems outperform expert humans in persuasive communication22 Jun 202681
- WIRED investigation: Flock Safety's AI person-search tools let police run broad description-based surveillance, with weak guardrails against misuse3 Sept 202680
- Hcompany open-sources NeoMME, a from-scratch multimodal-native encoder family, and NeoMME-Retriever for visual document retrieval3 Sept 202679
- Benchmark results show significant improvement in AI agent performance on WorkBench15 Jun 202679