Since sources are the foundation of Wikipedia’s verifiability, it is reasonable to ask: which websites play the most important role in its reference structure? To answer this question, references from tens of millions of Wikipedia articles across different languages were analysed.

The analysis was based on data current as of July 2026. The dataset included more than 400 million references (bibliographic citations) from approximately 68 million articles across more than 300 language editions of Wikipedia.

The ranking considers not only how frequently individual websites appear in Wikipedia references, but also the popularity of the articles in which those references occur. Article popularity is measured using the total number of page views over the previous 12 months.

Watch the full video:

Definition of a Website

In this analysis, a website means an identified web service, rather than merely an individual hostname or internet domain. For example, the services available at news.google.com and books.google.com are treated separately from the search engine available at google.com, even though all of them belong to the same registrable domain, google.com. As another example, ue.poznan.pl and poznan.pl are treated as separate websites.

Links to pages within the wikipedia.org domain were excluded from the analysis because, as a general rule, Wikipedia articles are not considered appropriate information sources for other Wikipedia articles.

Wikipedia Articles Included in the Calculations

Let 𝒜ref denote the set of Wikipedia articles containing at least one reference:

Aref={aARa>0}

Articles containing no references are not included in the formula. This ensures that the total number of references in every analysed article is greater than zero and that the denominator in the formula is always defined.

Pageview Weighting

Each website is evaluated using the PWRS metric (Pageview-Weighted Reference Score). For a website w, the score is calculated using the following formula:

PWRS(w)=aAref(PaRa,wRa)

where:

  • w – the website being analysed;
  • a – an individual Wikipedia article;
  • 𝒜 – the set of all Wikipedia articles included in the analysis;
  • 𝒜ref – the set of articles containing at least one reference;
  • Pa – the total number of page views received by article a over the previous 12 months;
  • Ra – the total number of references in article a;
  • Ra,w – the number of references in article a that point to website w;
  • PWRS(w) – the final Pageview-Weighted Reference Score assigned to website w.

Fair Popularity: Filtering and Adjustment of Pageview Data

The calculations included only pageviews classified in Wikimedia data as originating from users (agent = user), in accordance with Wikimedia’s definition of a pageview, which aims to distinguish user traffic from traffic generated by identified bots and web crawlers. However, this classification does not eliminate every case of anomalous traffic or traffic that does not reflect genuine reader interest. One example is the unusually high popularity of the Wikipedia article about Cleopatra, which has been linked to a default voice-search suggestion on devices using Google Assistant, as reported by Inverse. The Wikimedia Foundation also noted that the exceptionally high number of pageviews for this article did not reflect genuine reader interest. When compiling its list of the most popular articles of 2024, it therefore applied additional screening based, among other factors, on the share of views from desktop devices and the number of views without referrer information. For this reason, the present analysis employed additional procedures to identify anomalously high values and adjust them, usually by reducing the pageview count. This approach is conceptually similar to the fair popularity mechanism used by WikiRank since 2020. Consequently, the value Pa in the PWRS formula represents the corrected cumulative number of pageviews for article a classified as originating from users; the square-root transformation is then applied to this value.

Interpreting the Score

For each Wikipedia article, the number of references pointing to a given website is multiplied by the square root of that article’s page-view count and then divided by the total number of references in the article. The resulting values are then summed across all articles that meet the analysis criteria.

Taking pageviews into account gives greater weight to references appearing in more popular articles. At the same time, using the square root of the pageview count reduces the disproportionate influence of articles with exceptionally high traffic.

Dividing by the total number of references in an article also accounts for the visibility of individual sources. A reference appearing in an article with only a few citations receives greater weight than the same reference appearing in a very extensive reference list.

The resulting ranking measures the importance and visibility of websites within Wikipedia’s reference ecosystem under the proposed model. It should not be interpreted as a direct assessment of the truthfulness, reliability, or editorial quality of individual websites.

Top 25 Websites

PWRS values were calculated for individual websites in each language edition of Wikipedia. The following sections present rankings of the 25 websites with the highest PWRS values for each of the 28 selected language editions. The analysis included editions that, as of July 2026, contained at least 500,000 articles and had an article depth of at least 10 (current statistics are available at meta.wikimedia.org).

Multilingual Wikipedia (67.9M Articles)

  1. web.archive.org
  2. doi.org
  3. books.google.com
  4. worldcat.org
  5. nih.gov
  6. archive.org
  7. nytimes.com
  8. geonames.org
  9. catalogueoflife.org
  10. wikidata.org
  11. bbc.co.uk
  12. nasa.gov
  13. youtube.com
  14. adsabs.harvard.edu
  15. semanticscholar.org
  16. theguardian.com
  17. insee.fr
  18. census.gov
  19. deadline.com
  20. billboard.com
  21. variety.com
  22. imdb.com
  23. jstor.org
  24. newspapers.com
  25. webcitation.org

An extended version of the multilingual ranking is presented in this video.

English Wikipedia (7.2M Articles)

  1. web.archive.org
  2. doi.org
  3. books.google.com
  4. worldcat.org
  5. nih.gov
  6. archive.org
  7. nytimes.com
  8. semanticscholar.org
  9. adsabs.harvard.edu
  10. bbc.co.uk
  11. theguardian.com
  12. newspapers.com
  13. billboard.com
  14. youtube.com
  15. jstor.org
  16. deadline.com
  17. variety.com
  18. census.gov
  19. indiatimes.com
  20. allmusic.com
  21. latimes.com
  22. hollywoodreporter.com
  23. bbc.com
  24. washingtonpost.com
  25. espn.com

German Wikipedia (3.1M Articles)

  1. web.archive.org
  2. redirecter.toolforge.org
  3. doi.org
  4. books.google.de
  5. zdb-katalog.de
  6. spiegel.de
  7. nih.gov
  8. youtube.com
  9. imdb.com
  10. sueddeutsche.de
  11. welt.de
  12. faz.net
  13. archive.li
  14. zeit.de
  15. nytimes.com
  16. offiziellecharts.de
  17. orf.at
  18. filmdienst.de
  19. archive.today
  20. archive.org
  21. europa.eu
  22. theguardian.com
  23. digitale-sammlungen.de
  24. hitparade.ch
  25. bayern.de

French Wikipedia (2.8M Articles)

  1. insee.fr
  2. web.archive.org
  3. doi.org
  4. issn.org
  5. lemonde.fr
  6. bnf.fr
  7. wikiwix.com
  8. books.google.com
  9. books.google.fr
  10. lefigaro.fr
  11. nih.gov
  12. googleusercontent.com
  13. lequipe.fr
  14. culture.gouv.fr
  15. worldcat.org
  16. legifrance.gouv.fr
  17. imdb.com
  18. liberation.fr
  19. leparisien.fr
  20. allocine.fr
  21. youtube.com
  22. nasa.gov
  23. ouest-france.fr
  24. francetvinfo.fr
  25. nytimes.com

Swedish Wikipedia (2.6M Articles)

  1. catalogueoflife.org
  2. web.archive.org
  3. geonames.org
  4. kb.se
  5. nasa.gov
  6. iucnredlist.org
  7. ne.se
  8. smhi.se
  9. runeberg.org
  10. dn.se
  11. aftonbladet.se
  12. worldcat.org
  13. viewfinderpanoramas.org
  14. svt.se
  15. kit.edu
  16. doi.org
  17. dyntaxa.se
  18. www.istat.it
  19. scb.se
  20. svd.se
  21. expressen.se
  22. kew.org
  23. imdb.com
  24. sverigesradio.se
  25. archive.is

Dutch Wikipedia (2.2M Articles)

  1. web.archive.org
  2. marinespecies.org
  3. nos.nl
  4. nhm.ac.uk
  5. iucnredlist.org
  6. doi.org
  7. cbs.nl
  8. kb.nl
  9. ad.nl
  10. volkskrant.nl
  11. nu.nl
  12. nieuwsblad.be
  13. nrc.nl
  14. tamu.edu
  15. nih.gov
  16. archive.today
  17. hln.be
  18. youtube.com
  19. archive.org
  20. www.istat.it
  21. amnh.org
  22. standaard.be
  23. vrt.be
  24. trouw.nl
  25. gbif.fr

Spanish Wikipedia (2.1M Articles)

  1. web.archive.org
  2. doi.org
  3. issn.org
  4. archive.org
  5. nih.gov
  6. elpais.com
  7. books.google.com
  8. youtube.com
  9. books.google.es
  10. nytimes.com
  11. archive.today
  12. infobae.com
  13. worldcat.org
  14. deadline.com
  15. imdb.com
  16. elmundo.es
  17. worldpostalcodes.org
  18. abc.es
  19. bbc.co.uk
  20. billboard.com
  21. variety.com
  22. lanacion.com.ar
  23. census.gov
  24. marca.com
  25. theguardian.com

Russian Wikipedia (2.1M Articles)

  1. web.archive.org
  2. webcitation.org
  3. worldcat.org
  4. gks.ru
  5. doi.org
  6. archive.org
  7. wikidata.org
  8. ru.wikisource.org
  9. archive.today
  10. youtube.com
  11. nih.gov
  12. books.google.com
  13. kommersant.ru
  14. rosstat.gov.ru
  15. nytimes.com
  16. ria.ru
  17. books.google.ru
  18. deadline.com
  19. imdb.com
  20. lenta.ru
  21. tass.ru
  22. narod.ru
  23. bbc.com
  24. variety.com
  25. bbc.co.uk

Italian Wikipedia (2.0M Articles)

  1. web.archive.org
  2. archive.org
  3. doi.org
  4. repubblica.it
  5. books.google.it
  6. nih.gov
  7. treccani.it
  8. gazzetta.it
  9. corriere.it
  10. books.google.com
  11. archive.is
  12. youtube.com
  13. www.istat.it
  14. oadoi.org
  15. nytimes.com
  16. insee.fr
  17. allmusic.com
  18. worldcat.org
  19. deadline.com
  20. imdb.com
  21. antoniogenna.net
  22. ansa.it
  23. variety.com
  24. sky.it
  25. sbn.it

Polish Wikipedia (1.7M Articles)

  1. web.archive.org
  2. worldcat.org
  3. sejm.gov.pl
  4. archive.is
  5. doi.org
  6. stat.gov.pl
  7. nih.gov
  8. pwn.pl
  9. dane.gov.pl
  10. wyborcza.pl
  11. archive.ph
  12. onet.pl
  13. wp.pl
  14. discogs.com
  15. wirtualnemedia.pl
  16. youtube.com
  17. imdb.com
  18. allmusic.com
  19. filmpolski.pl
  20. pkw.gov.pl
  21. nauka-polska.pl
  22. interia.pl
  23. books.google.pl
  24. archive.org
  25. filmweb.pl

Chinese Wikipedia (1.5M Articles)

  1. web.archive.org
  2. archive.org
  3. nih.gov
  4. worldcat.org
  5. doi.org
  6. books.google.com
  7. sina.com.cn
  8. youtube.com
  9. archive.today
  10. ltn.com.tw
  11. natalie.mu
  12. hk01.com
  13. census.gov
  14. qq.com
  15. sohu.com
  16. zh.wikisource.org
  17. people.com.cn
  18. nationalmap.gov
  19. xinhuanet.com
  20. udn.com
  21. nytimes.com
  22. webcitation.org
  23. geoportail.gouv.fr
  24. cna.com.tw
  25. mca.gov.cn

Japanese Wikipedia (1.5M Articles)

  1. web.archive.org
  2. natalie.mu
  3. doi.org
  4. oricon.co.jp
  5. ndl.go.jp
  6. x.com
  7. nikkansports.com
  8. youtube.com
  9. sponichi.co.jp
  10. asahi.com
  11. nih.gov
  12. worldcat.org
  13. nikkei.com
  14. kotobank.jp
  15. nii.ac.jp
  16. nhk.or.jp
  17. sankei.com
  18. twitter.com
  19. mainichi.jp
  20. eiga.com
  21. prtimes.jp
  22. sanspo.com
  23. ameblo.jp
  24. realsound.jp
  25. yomiuri.co.jp

Ukrainian Wikipedia (1.4M Articles)

  1. web.archive.org
  2. webcitation.org
  3. doi.org
  4. rada.gov.ua
  5. nih.gov
  6. worldcat.org
  7. stat.gov.pl
  8. insee.fr
  9. archive.org
  10. census.gov
  11. youtube.com
  12. books.google.com
  13. president.gov.ua
  14. ukrcensus.gov.ua
  15. pravda.com.ua
  16. archive.today
  17. deadline.com
  18. suspilne.media
  19. rbc.ua
  20. imdb.com
  21. history.org.ua
  22. gks.ru
  23. bbc.com
  24. nytimes.com
  25. loc.gov

Arabic Wikipedia (1.3M Articles)

  1. web.archive.org
  2. wikidata.org
  3. archive.org
  4. doi.org
  5. worldcat.org
  6. nih.gov
  7. books.google.com
  8. issn.org
  9. openlibrary.org
  10. elcinema.com
  11. imdb.com
  12. geonames.org
  13. census.gov
  14. nytimes.com
  15. bbc.co.uk
  16. shamela.ws
  17. minorplanetcenter.net
  18. aljazeera.net
  19. loc.gov
  20. theguardian.com
  21. bbc.com
  22. britannica.com
  23. webcitation.org
  24. youtube.com
  25. wikidata-externalid-url.toolforge.org

Vietnamese Wikipedia (1.3M Articles)

  1. web.archive.org
  2. theplantlist.org
  3. doi.org
  4. archive.org
  5. catalogueoflife.org
  6. books.google.com
  7. tamu.edu
  8. nih.gov
  9. worldcat.org
  10. vnexpress.net
  11. tuoitre.vn
  12. archive.today
  13. bbc.co.uk
  14. marinespecies.org
  15. thanhnien.vn
  16. chinhphu.vn
  17. amnh.org
  18. thuvienphapluat.vn
  19. census.gov
  20. vietnamnet.vn
  21. nytimes.com
  22. dantri.com.vn
  23. nhm.ac.uk
  24. youtube.com
  25. adsabs.harvard.edu

Portuguese Wikipedia (1.2M Articles)

  1. web.archive.org
  2. globo.com
  3. uol.com.br
  4. doi.org
  5. worldcat.org
  6. nih.gov
  7. books.google.com
  8. ibge.gov.br
  9. insee.fr
  10. books.google.com.br
  11. nytimes.com
  12. terra.com.br
  13. archive.org
  14. abril.com.br
  15. www.istat.it
  16. estadao.com.br
  17. bbc.co.uk
  18. census.gov
  19. billboard.com
  20. webcitation.org
  21. archive.is
  22. deadline.com
  23. theguardian.com
  24. sapo.pt
  25. adsabs.harvard.edu

Persian Wikipedia (1.1M Articles)

  1. web.archive.org
  2. doi.org
  3. nih.gov
  4. archive.org
  5. books.google.com
  6. bbc.com
  7. worldcat.org
  8. isna.ir
  9. bbc.co.uk
  10. mehrnews.com
  11. irna.ir
  12. amar.org.ir
  13. nytimes.com
  14. radiofarda.com
  15. webcitation.org
  16. archive.today
  17. iranicaonline.org
  18. hamshahrionline.ir
  19. sci.org.ir
  20. iranshahrpedia.com
  21. khabaronline.ir
  22. theguardian.com
  23. deadline.com
  24. boxofficemojo.com
  25. tasnimnews.com

Chechen Wikipedia (866K Articles)

  1. web.archive.org
  2. gmcrosstata.ru
  3. worldcat.org
  4. wikidata.org
  5. insee.fr
  6. tuik.gov.tr
  7. nasa.gov
  8. czso.cz
  9. census.gov
  10. geonames.org
  11. trasa.ru
  12. ine.es
  13. www.istat.it
  14. e-local.gob.mx
  15. ibge.gov.br
  16. minorplanetcenter.net
  17. doi.org
  18. mind.ua
  19. stat.gov.pl
  20. webcitation.org
  21. gks.ru
  22. amar.org.ir
  23. nsi.bg
  24. inegi.org.mx
  25. unam.mx

Catalan Wikipedia (796K Articles)

  1. web.archive.org
  2. insee.fr
  3. books.google.cat
  4. doi.org
  5. gencat.cat
  6. worldcat.org
  7. nih.gov
  8. enciclopedia.cat
  9. census.gov
  10. archive.today
  11. elpais.com
  12. fishbase.org
  13. lavanguardia.com
  14. nytimes.com
  15. archive.org
  16. ara.cat
  17. termcat.cat
  18. ccma.cat
  19. vilaweb.cat
  20. britannica.com
  21. bbc.co.uk
  22. adsabs.harvard.edu
  23. imdb.com
  24. iec.cat
  25. esadir.cat

Indonesian Wikipedia (783K Articles)

  1. web.archive.org
  2. doi.org
  3. archive.org
  4. books.google.com
  5. worldcat.org
  6. nih.gov
  7. kompas.com
  8. detik.com
  9. bps.go.id
  10. books.google.co.id
  11. tribunnews.com
  12. liputan6.com
  13. tempo.co
  14. kemdikbud.go.id
  15. naver.com
  16. bbc.co.uk
  17. antaranews.com
  18. youtube.com
  19. adsabs.harvard.edu
  20. nytimes.com
  21. kemendagri.go.id
  22. kapanlagi.com
  23. cnnindonesia.com
  24. archive.today
  25. google.co.id

Korean Wikipedia (752K Articles)

  1. web.archive.org
  2. naver.com
  3. doi.org
  4. nih.gov
  5. archive.org
  6. worldcat.org
  7. books.google.com
  8. adsabs.harvard.edu
  9. chosun.com
  10. law.go.kr
  11. semanticscholar.org
  12. yna.co.kr
  13. daum.net
  14. ko.wikisource.org
  15. donga.com
  16. hani.co.kr
  17. nytimes.com
  18. youtube.com
  19. archive.today
  20. khan.co.kr
  21. history.go.kr
  22. aks.ac.kr
  23. census.gov
  24. bbc.co.uk
  25. mt.co.kr

Serbian Wikipedia (706K Articles)

  1. inegi.org.mx
  2. web.archive.org
  3. geonames.org
  4. stat.gov.rs
  5. doi.org
  6. missouri.edu
  7. nih.gov
  8. archive.org
  9. www.istat.it
  10. books.google.com
  11. worldcat.org
  12. rts.rs
  13. webcitation.org
  14. novosti.rs
  15. iucnredlist.org
  16. insee.fr
  17. politika.rs
  18. blic.rs
  19. kia.hu
  20. u-strasbg.fr
  21. b92.net
  22. britannica.com
  23. nb.rs
  24. dzs.hr
  25. youtube.com

Turkish Wikipedia (689K Articles)

  1. web.archive.org
  2. doi.org
  3. tuik.gov.tr
  4. archive.org
  5. books.google.com
  6. worldcat.org
  7. hurriyet.com.tr
  8. nih.gov
  9. mevzuat.gov.tr
  10. webcitation.org
  11. milliyet.com.tr
  12. itis.gov
  13. gbif.org
  14. archive.today
  15. youtube.com
  16. yerelnet.org.tr
  17. tbmm.gov.tr
  18. bbc.co.uk
  19. nytimes.com
  20. dergipark.org.tr
  21. bbc.com
  22. sabah.com.tr
  23. aa.com.tr
  24. haberturk.com
  25. books.google.com.tr

Norwegian Wikipedia (Bokmål) (686K Articles)

  1. web.archive.org
  2. nb.no
  3. wikidata.org
  4. nrk.no
  5. snl.no
  6. doi.org
  7. vg.no
  8. census.gov
  9. worldcat.org
  10. wikidata-externalid-url.toolforge.org
  11. aftenposten.no
  12. imdb.com
  13. dagbladet.no
  14. ssb.no
  15. artsdatabanken.no
  16. olympics.com
  17. britannica.com
  18. nve.no
  19. olympedia.org
  20. transfermarkt.com
  21. regjeringen.no
  22. tv2.no
  23. sports-reference.com
  24. bnf.fr
  25. nih.gov

Urdu Wikipedia (641K Articles)

  1. web.archive.org
  2. nasa.gov
  3. espncricinfo.com
  4. books.google.com
  5. archive.org
  6. imdb.com
  7. cricketarchive.com
  8. minorplanetcenter.net
  9. d-nb.info
  10. geonames.org
  11. doi.org
  12. bnf.fr
  13. worldcat.org
  14. dawn.com
  15. census.gov
  16. nkp.cz
  17. openstreetmap.org
  18. indiatimes.com
  19. britannica.com
  20. snaccooperative.org
  21. wikidata.org
  22. nih.gov
  23. bbc.co.uk
  24. findagrave.com
  25. shamela.ws

Finnish Wikipedia (621K Articles)

  1. web.archive.org
  2. yle.fi
  3. hs.fi
  4. britannica.com
  5. is.fi
  6. iltalehti.fi
  7. maanmittauslaitos.fi
  8. wikidata.org
  9. worldcat.org
  10. eliteprospects.com
  11. doi.org
  12. iucnredlist.org
  13. nih.gov
  14. ifpi.fi
  15. citypopulation.de
  16. mtvuutiset.fi
  17. kansalliskirjasto.fi
  18. books.google.fi
  19. allmusic.com
  20. syke.fi
  21. kansallisbiografia.fi
  22. bbc.co.uk
  23. bbc.com
  24. nytimes.com
  25. helsinki.fi

Czech Wikipedia (594K Articles)

  1. web.archive.org
  2. worldcat.org
  3. idnes.cz
  4. doi.org
  5. archive.org
  6. csu.gov.cz
  7. ceskatelevize.cz
  8. nih.gov
  9. novinky.cz
  10. czso.cz
  11. denik.cz
  12. rozhlas.cz
  13. cas.cz
  14. volby.cz
  15. aktualne.cz
  16. psp.cz
  17. lidovky.cz
  18. irozhlas.cz
  19. pamatkovykatalog.cz
  20. seznamzpravy.cz
  21. youtube.com
  22. books.google.cz
  23. books.google.com
  24. iucnredlist.org
  25. muni.cz

Hungarian Wikipedia (571K Articles)

  1. web.archive.org
  2. valasztas.hu
  3. doi.org
  4. oszk.hu
  5. archive.org
  6. nasa.gov
  7. iszdb.hu
  8. nih.gov
  9. nemzetisport.hu
  10. index.hu
  11. www.istat.it
  12. ksh.hu
  13. familysearch.org
  14. destatis.de
  15. origo.hu
  16. youtube.com
  17. imdb.com
  18. worldcat.org
  19. census.gov
  20. hvg.hu
  21. arcanum.com
  22. pim.hu
  23. 24.hu
  24. bnf.fr
  25. bbc.co.uk

Romanian Wikipedia (545K Articles)

  1. web.archive.org
  2. wikidata.org
  3. doi.org
  4. toutes-les-villes.com
  5. citypopulation.de
  6. nih.gov
  7. europa.eu
  8. books.google.com
  9. worldcat.org
  10. nasa.gov
  11. minorplanetcenter.net
  12. ukrcensus.gov.ua
  13. geopf.fr
  14. www.istat.it
  15. catalogueoflife.org
  16. bnf.fr
  17. adevarul.ro
  18. recensamantromania.ro
  19. archive.org
  20. imdb.com
  21. destatis.de
  22. d-nb.info
  23. wikidata-externalid-url.toolforge.org
  24. nytimes.com
  25. bbc.co.uk

Discussion and Limitations

Interpretation of the Term “Information Source”

The presence of a particular website address in a Wikipedia reference does not always mean that the website is the original creator or publisher of the information. In some cases, it primarily serves a technical, intermediary, archival, or discovery function. One example is doi.org, the infrastructure used to resolve persistent DOI identifiers. A DOI address identifies a specific object, such as a scholarly publication, and then redirects the user to the object’s current location or description. The doi.org website therefore often acts as an intermediary leading to a publisher’s website, repository, or database rather than serving as the actual publisher of the cited work. Determining which websites individual DOI identifiers ultimately lead to would require resolving the identifiers and tracing redirects at the level of individual references, which could be the subject of a separate study. More information about how DOIs work is available on the DOI Foundation website.

A similar function is performed by hdl.handle.net, which acts as a proxy server for the Handle System. It reads a persistent identifier, locates the address assigned to it, and redirects the user to the target resource. In this case, the number of references attributed to hdl.handle.net reflects the importance of the identifier infrastructure rather than that of a specific content provider. The operation of this mechanism is described in the Handle.Net documentation.

An even more technical example is redirecter.toolforge.org. This service accepts a target address in the url parameter and redirects the user to the specified location. It is therefore not a source of content itself, but an intermediary tool. Identifying the actual website would require reading the redirect parameter or tracing the server response. The tool’s function can be checked directly at redirecter.toolforge.org.

Web Archiving Services

Web archives such as web.archive.org, webcitation.org, and ghostarchive.org form a separate category. They are generally not the original publishers of the cited information; instead, they store copies of pages from other websites and provide access to their content as it appeared at a particular point in time. The Wayback Machine (web.archive.org) allows users to browse preserved versions of web pages from different periods. An archive URL therefore points both to the archiving service and to the original website from which the preserved content came. WebCite and Ghostarchive serve a similar role by storing copies of pages to ensure later access to content that may be changed or removed.

Classifying an archive as a source is not necessarily an error. For historical information, a preserved copy of a page may be the only available basis for verifying a claim. Nevertheless, the website providing the archived copy should be distinguished from the website that originally published the content.

Catalogues and Access Platforms

Interpretive uncertainty also arises with bibliographic catalogues and publication access platforms. For example, worldcat.org is a global catalogue of library collections that helps users find a book or other document and determine which libraries hold it. Depending on the context, the appropriate information source may be either the WorldCat bibliographic record or the publication it describes. WorldCat’s role as a catalogue and resource-discovery tool is explained by OCLC.

A similar issue arises with books.google.com. Google Books allows users to search, browse, and, in some cases, read scanned books, but the source of the content is generally the specific book together with its author and publisher. Google Books may therefore function simultaneously as an access platform, a search engine, a repository of digital copies, and a source of bibliographic data. The scope of its functions is described in the Google Books documentation.

Selection of Links Included in the Ranking

The examples above show that deciding which links to include in the ranking is a methodological choice. One approach is to retain all websites that appear directly in reference URLs; another is to exclude services that are purely technical; a third is to attribute links to the websites ultimately reached through redirects. Each approach answers a somewhat different research question.

Including intermediary services makes it possible to measure their importance within Wikipedia’s reference infrastructure. Excluding them, or reassigning references to the target websites, would provide a better account of the origin of the content. However, doing so would require additional analysis of URLs, redirect parameters, HTTP response chains, persistent identifiers, and original addresses embedded in archive links.

Some websites may also perform several functions at once. A bibliographic catalogue may be a source of information about a book edition, an archive may be a source of information about the historical state of a page, and a repository may both store a document and provide its metadata. Automatically dividing all websites into “sources” and “intermediaries” may therefore be impossible without analysing the context of each individual reference.

Limitations of the PWRS Metric

PWRS is only one possible way to measure the importance of websites in Wikipedia references. The score describes a website’s importance under the adopted model, but it is not a direct measure of reliability, scholarly quality, social significance, or the informational value of the published content. An article’s popularity is not equivalent to its scholarly or encyclopaedic importance. Articles about entertainment, popular culture, sport, or current events may receive far more page views than articles covering topics fundamental to science, history, or education. As a result, websites associated with subjects that are particularly popular among the readers of a given language edition may rank higher. For this reason, the PWRS formula uses the square root of the page-view count. This transformation reduces the differences between articles of moderate and exceptionally high popularity while still assigning greater weight to references appearing in more frequently read articles. The square root limits the influence of extremely popular articles but does not eliminate it entirely.

The use of the square root is one possible modelling choice. Alternative approaches could use, among other methods, the logarithm of the page-view count, capping of extreme values, a percentile scale, equal weighting of all articles, or normalisation by subject category. Other measures could also be incorporated, such as the number of distinct articles referring to a website, the number of language editions in which the website appears, link persistence, or the thematic diversity of the citing articles.

Further research could therefore include a sensitivity analysis of the ranking under different page-view weighting methods, a classification of websites by function, and the reassignment of intermediary and archival links to the original content providers. This would make it possible to compare the importance of websites as direct information sources with their importance as components of the technical infrastructure used to access knowledge.

A broader comparison of different methods for assessing source importance is available through BestRef, which uses several computational models. The service provides scores for several million websites, calculated separately for each of more than 300 Wikipedia language editions.

More information about the analysis of Wikipedia information sources can be found in the following scholarly publications: