Cross-population polygenic risk scores: when European data stops helping
Most genetic risk models are built from European cohorts and lose accuracy elsewhere. A Google Research study of eight traits in Biobank Japan tests how cross-population polygenic risk scores should mix European and local data, and finds that the answer shifts with sample size and with the trait.

Why do polygenic risk scores lose accuracy outside Europe?
Polygenic risk scores add up the small effects of many genetic variants, often hundreds to millions, into one estimate of disease risk or of a measured trait. Most of the association studies behind them were run on people of European ancestry, so the scores travel poorly. Genetic architecture, population structure and allele frequencies all differ between populations, and each difference erodes polygenic risk score accuracy when a European-trained score is applied elsewhere.
That gap limits clinical use. Martin and colleagues reported in 2019 that the scores then available were several times more accurate for people of European descent than for other groups, and warned that using them in clinics could widen health disparities. The Google Research post adds a practical constraint: running a fresh genome-wide association study (GWAS) across hundreds of thousands of people costs more than most health systems can spend.
Transfer learning for genomic prediction is the natural candidate. Start from large European GWAS results and supplement them with data from the target population. The question the study tackles is how much of each to use, and at what point the European data stops adding accuracy.
How did the study test cross-population polygenic risk scores?
The post's authors, Joey Poomarin Phloyphisut and Cory McLean of Google Research, describe a series of controlled dataset ablation experiments on two large, deeply phenotyped cohorts. European participants in UK Biobank served as the source population. Biobank Japan, a cohort of nearly 200 thousand Japanese individuals, served as the target. Because both cohorts record the same traits, the team could subsample either one on purpose and watch how accuracy responded.
They chose eight clinically relevant traits recorded in both cohorts: body mass index (BMI), systolic and diastolic blood pressure, red and white blood cell counts, HDL and LDL cholesterol, and blood glucose. In UK Biobank, the share of trait variation explained by the measured variants (the SNP heritability) ranged from 0.07 to 0.28.
Every model was scored on one fixed, held-out set of Biobank Japan participants, using the Pearson correlation between predicted and measured values. The stated goal is practical: evidence-based guidance on how cross-population GWAS and risk-score training should be set up to get the best accuracy in the target population.
Three ways to combine European and Japanese data
The study compared three pipelines. They differ in where Japanese data enters: only at model training, also at variant discovery, or through a method designed to weigh several ancestries at once.
- UK Biobank discovery with elastic net. A GWAS on the full European UK Biobank cohort picked the candidate variants, keeping only independent variants that also exist in the Japanese data. Elastic net models, a penalised form of linear regression, were then trained on many mixes of Japanese and European samples: 12 to 13 Biobank Japan sizes crossed with seven UK Biobank sizes; the post reports 96 to 104 models per trait.
- Meta-analysis with elastic net. Separate GWAS were run on the full UK Biobank cohort and on a sampled Japanese subset. Their results were merged by meta-analysis to choose the variants, and elastic net models were then trained on mixed European and Japanese data.
- PRS-CSx. PRS-CSx is a published method that links genetic effects across ancestries through a shared shrinkage prior and accounts for differences in linkage disequilibrium between populations. It was fitted to the full European data and to sampled Japanese data. Because it returns one score per population, part of the validation set was used to find the best linear blend of the two.
When does more European data stop helping?
Once the Japanese training set is no longer small. The authors report that European data gives a useful baseline when target-population data is limited or absent, but that models trained on Japanese data alone pull ahead from about 15 thousand Japanese samples. Beyond that point, co-training with European data appears to hold back the gains that more local samples would otherwise bring.
HDL cholesterol gives the clearest example. Past a crossover at 15,000 Biobank Japan samples, adding more than 5k UK Biobank samples lowered accuracy compared with training on Japanese data only. At the other extreme, with a very small target set such as 5,000 samples, pooling European data gave a clear statistical lift. The authors say the broad pattern, where out-of-population data can actively hurt, held for every trait they examined.
Analysis: in machine learning terms this is negative transfer. The European samples behave like a strong prior. When local data is thin, the prior fills the gaps; when local data is plentiful, it pulls the model towards effect sizes that fit the target population less well.
Shared genetics sets how long European data stays useful
The crossover depends on how much of a trait's genetics the two populations share. The authors measure this with genetic correlation: how closely a trait's genetic effects agree across populations. Traits with higher genetic correlation, which they call conserved, kept gaining from pooled European data up to roughly 25 to 40 thousand or more Japanese samples before Japanese-only training caught up.
Five of the eight traits, BMI among them, fall in this shared group. For these traits, the best amount of European data stays at the largest size tested until the Japanese set grows past about 40 thousand, and only then starts to fall.
Lipid levels (HDL and LDL) and blood glucose behave differently. Their crossover arrives at much smaller Japanese sample sizes, and the best amount of European data is smaller as well. The authors attribute this to the European data sitting further from the target distribution for these more population-specific traits.
How did meta-analysis and PRS-CSx compare?
Each helped in a different regime. The main experiments used only variants discovered in European data, so variants linked to a trait solely in Japanese participants were missed. Meta-analysis and PRS-CSx both bring Japanese GWAS results into the picture, at the cost of relying on much smaller Japanese discovery samples.
For conserved traits, meta-analysis made little difference, largely because the smaller Japanese GWAS lacked statistical power. Population-specific traits told another story: for HDL and LDL, and to a lesser degree blood glucose, meta-analysis clearly beat discovery in a single population. The authors trace most of that gain to the changed set of input variants. Adding European samples at training time gave a slight further improvement only with 10,000 or fewer Japanese samples.
PRS-CSx weights population-specific models dynamically, so in theory it should cope with both kinds of trait. In practice the authors found it needed more data than elastic net. Below 25 thousand target samples it trailed the strongest elastic net model for every trait except BMI. As target samples approached 100 thousand, it matched or beat the best model for every trait except blood glucose.
Limitations and open questions
The findings come from one source population, one target population and eight continuous traits, and they are reported in a blog post. As of 1 October 2026, the post links no paper, preprint or code, so details such as per-trait sample sizes and exact correlation values can be checked only against the post's own figures.
- Population scope. Only European-to-Japanese transfer was tested. Analysis: crossover points for other target ancestries, or for several source cohorts used together, may differ and would need their own titration experiments.
- Outcome scope. All eight traits are quantitative measurements scored by Pearson correlation. Analysis: the post does not report binary disease outcomes, calibration or clinical decision metrics, which would matter for any real deployment.
- Variant discovery. The main experiments restricted inputs to variants found in UK Biobank, which the authors acknowledge leaves out variants unique to Japanese participants. Meta-analysis and PRS-CSx address this only in part, because the Japanese GWAS are small.
- The authors' conclusion. Better accuracy in diverse populations will need both larger local, diverse biobanks and modelling choices matched to each trait's heritability and to the sample size available.
What does this mean for applied and enterprise AI?
Analysis: the study is a domain-specific but unusually clean measurement of when transfer learning helps and when it hurts. The lesson carries over to any team that trains on a large source dataset and adapts to a smaller, different target, such as a new market, region or customer segment.
- Measure the crossover. The authors' titration grid, crossing several target sizes with several source sizes, is a reusable design. It shows where extra source data starts to cost accuracy instead of assuming that more data is always better.
- Use similarity to predict transfer. Genetic correlation told the authors how long European data stayed useful. Teams can look for an equivalent measure of source and target similarity before deciding how much external data to mix in.
- Simple methods can win at small sizes. PRS-CSx was built for multi-ancestry data yet needed far more samples than elastic net to perform well. A more specialised model is not automatically the better choice when in-domain data is scarce.
- Evaluate on the target. Every model in the study was judged on the same held-out Japanese set. Judging transfer by performance on the source population would not reveal the effect.
Data, code and model availability
The post does not announce a code, model or data release. Both cohorts are established research resources run by their own organisations: UK Biobank holds genetic and phenotypic data on approximately 500,000 people across the United Kingdom, and Biobank Japan registered 200,000 participants in its first five-year enrolment period.
PRS-CSx is described in Ruan and colleagues' 2022 paper, listed in the references. For this specific study, the post remains the only public description so far.
Questions and answers
What is a cross-population polygenic risk score?
It is a polygenic risk score built partly from genetic data on one population and applied to another. In this study, European data from UK Biobank was used to help predict eight traits in Japanese participants from Biobank Japan. The aim is to borrow statistical strength from large, mostly European studies when the target population has far fewer genotyped and measured people.
Does adding European data always improve genetic risk prediction for other populations?
No. The authors report that European UK Biobank data improved prediction for Japanese participants only while the Japanese training set was small. Past roughly 15,000 Biobank Japan samples, models trained on Japanese data alone did better, and for HDL cholesterol, adding more than 5k European samples beyond that point reduced accuracy. The exact crossover varied from trait to trait.
Which traits benefited most from European data?
Traits whose genetic effects are highly correlated between the two populations, such as BMI, kept benefiting from pooled European data up to roughly 25 to 40 thousand Japanese samples or more. Lipid traits (HDL and LDL cholesterol) and blood glucose, which have more population-specific genetics, lost the benefit at much smaller Japanese sample sizes.
When did PRS-CSx outperform simpler methods?
Only with plenty of target data. The authors found that PRS-CSx performed worse than the strongest elastic net model for every trait except BMI when the Japanese sample was under 25 thousand. As the Japanese sample approached 100 thousand, it matched or exceeded the best model for every trait except blood glucose.
What can machine learning teams outside genomics learn from this study?
Analysis: transfer from a large but mismatched source dataset has a crossover point. Below it, source data fills gaps; above it, the same data can lower accuracy on the target. Measuring performance across a grid of source and target sizes, on held-out target data, shows where that point lies instead of assuming that more data always helps.
References
- Ruan, Y., Lin, Y. F., Feng, Y. A., Chen, C. Y., Lam, M., Guo, Z., Stanley Global Asia Initiatives, He, L., Sawa, A., Martin, A. R., Qin, S., Huang, H., & Ge, T. (2022). Improving polygenic prediction in ancestrally diverse populations. Nature Genetics, 54(5), 573-580. https://doi.org/10.1038/s41588-022-01054-7 (external site)
- Martin, A. R., Kanai, M., Kamatani, Y., Okada, Y., Neale, B. M., & Daly, M. J. (2019). Clinical use of current polygenic risk scores may exacerbate health disparities. Nature Genetics, 51(4), 584-591. https://doi.org/10.1038/s41588-019-0379-x (external site)
- Bycroft, C., Freeman, C., Petkova, D., Band, G., Elliott, L. T., Sharp, K., Motyer, A., Vukcevic, D., Delaneau, O., O'Connell, J., Cortes, A., Welsh, S., Young, A., Effingham, M., McVean, G., Leslie, S., Allen, N., Donnelly, P., & Marchini, J. (2018). The UK Biobank resource with deep phenotyping and genomic data. Nature, 562(7726), 203-209. https://doi.org/10.1038/s41586-018-0579-z (external site)
- Nagai, A., Hirata, M., Kamatani, Y., Muto, K., Matsuda, K., Kiyohara, Y., Ninomiya, T., Tamakoshi, A., Yamagata, Z., Mushiroda, T., Murakami, Y., Yuji, K., Furukawa, Y., Zembutsu, H., Tanaka, T., Ohnishi, Y., Nakamura, Y., BioBank Japan Cooperative Hospital Group, & Kubo, M. (2017). Overview of the BioBank Japan Project: Study design and profile. Journal of Epidemiology, 27(3S), S2-S8. https://doi.org/10.1016/j.je.2016.12.005 (external site)
Original article
Phloyphisut, J. P., & McLean, C. (2026, 3 September). Transfer learning for genomic prediction in underrepresented populations. Google Research Blog. https://research.google/blog/transfer-learning-for-genomic-prediction-in-underrepresented-populations/ (external site)