A Python library that provides functions to retrieve names, ISO and FIPS codes of continents, countries and first- and second-level administrative divisions as well as US states and counties as Python dictionaries. The country and city datasets also include population and geographic data.
Geonames data is obtained from GeoNames.
pip install geonamescache
A simple example:
import geonamescache
gc = geonamescache.GeonamesCache()
print(gc.get_countries())
The datasets are bundled gzipped and parsed on first use, so the installed package is about 25 MB. Each GeonamesCache instance caches every dataset it loads, so keep one instance around rather than creating a new one per lookup.
When creating a GeonamesCache you can set the min_city_population parameter to either of 500, 1000, 5000 or the default 15000. The smaller the minimum population the more cities are included in the cities dataset.
Currently geonamescache provides the following methods, that return dictionaries with the requested data:
- get_continents()
- get_countries()
- get_admin1_codes()
- get_admin2_codes()
- get_us_states()
- get_cities()
- get_countries_by_names()
- get_us_states_by_names()
- get_cities_by_name(name)
- get_cities_by_names()
- get_us_counties()
- get_timezones()
In addition you can search for cities and administrative divisions by name.
- search_cities('NAME', countrycode=None, admin1code=None, case_sensitive=False, contains_search=True)
- search_admin1('NAME', countrycode=None, case_sensitive=False, contains_search=False, historic=False)
- search_admin2('NAME', countrycode=None, admin1code=None, case_sensitive=False, contains_search=False, historic=False)
Each returns a flat list of records matching NAME.
countrycodeandadmin1coderestrict the search before names are compared. Place names are not unique — "Santa Rosa" names eight second-level divisions worldwide — so an unscoped search is rarely the answer you want.admin1codetakes the bare code (11) or the composite one (CO.11), and is ignored without acountrycode.- By default the city search looks at the
alternatenamesattribute; passattributefor another one. The division searches always covername,asciiname,englishnameand every alternate name. - An exact, case insensitive
search_cities()is answered from an index of every value keyed by country and casefolded value, built once per attribute on first use, so repeated lookups cost a dict access instead of a pass over the dataset. The other combinations still scan, but acountrycodenarrows the scan to that country's records. Indexed results come out in country order rather than geonameid order. - By default the search is case insensitive, it can be made case sensitive by changing
case_sensitiveto True. - The city search is a contains search by default; the division searches are exact by default, because a substring of a word as common as "north" matches hundreds of divisions. Either can be switched with
contains_search. historic=Truealso matches names the source marks as superseded — Venezuela'sVE.26still answers to "Vargas", renamed La Guaira in 2019. Off by default, because a historic name can now belong somewhere else.
To get a country's, a division's or a city's names in every language, or a few of them:
- get_country_names(country, languages=None)
- get_admin1_names(admin1, languages=None, historic=False)
- get_city_names(city, languages=None, historic=False)
To resolve the administrative division a city belongs to, use:
- get_admin1_by_city(city)
- get_admin2_by_city(city)
Both take a city record and return the matching division record, or None if the city's codes are missing or not present in the division dataset.
To get the time zones of a country, use:
- get_timezones_by_country(countrycode)
All examples below assume gc = geonamescache.GeonamesCache().
A dictionary keyed by the two-letter continent code. Records come from the GeoNames web service and contain more fields than shown here.
>>> gc.get_continents()['EU']['name']
'Europe'
A dictionary of 252 countries keyed by ISO alpha-2 code.
>>> gc.get_countries()['US']
{
'geonameid': 6252001,
'name': 'United States',
'iso': 'US',
'iso3': 'USA',
'isonumeric': 840,
'fips': 'US',
'continentcode': 'NA',
'capital': 'Washington',
'areakm2': 9629091,
'population': 327167434,
'tld': '.us',
'currencycode': 'USD',
'currencyname': 'Dollar',
'phone': '1',
'postalcoderegex': '^\\d{5}(-\\d{4})?$',
'languages': 'en-US,es-US,haw,fr',
'neighbours': 'CA,MX,CU',
'alternatenames': {
'en': ['United States of America', 'United States', 'America', 'USA', ...],
'fr': ['États-Unis', ...],
...
}
}
get_countries_by_names() returns the same records keyed by country name instead, e. g. gc.get_countries_by_names()['Spain'].
alternatenames holds the country's other names grouped by ISO-639 language code, with the preferred name of each language first:
>>> gc.get_countries()['NL']['alternatenames']['fr']
['Pays-Bas']
They are grouped rather than flat — unlike city alternate names, whose source column records no language — because there are a great many of them: 43,600 across all countries, and 163 languages for the United Kingdom alone. A caller matching text in a few languages can then take only those. Historic names and rows that are reference codes rather than names (link, wkdt, post, iata, icao, faac, unlc, tcid, phon, piny) are excluded when the data is built.
A country's name plus its alternate names, as one deduplicated list with the name first. This is the form to use when matching names found in text:
>>> gc.get_country_names(gc.get_countries()['GB'])[:4]
['United Kingdom', 'Great Britain', 'Britain', 'UK']
Pass languages to take only some of them. Note that 163 languages of names is a lot to match against, so narrowing to the languages you actually expect is usually worth it:
>>> gc.get_country_names(gc.get_countries()['GB'], languages=('pt', 'es'))
['United Kingdom', 'Britain', 'Reino Unido', 'RU']
("Britain" is there because it is one of the unlanguaged names described below, not because it is Portuguese or Spanish.)
Names recorded without a language are always included, whatever languages says, because they are language-agnostic and some are the form most used in English. The Netherlands is the example that matters: its name is "The Netherlands" and its en names do not contain the bare "Netherlands", which sits under the empty-language key.
>>> 'Netherlands' in gc.get_country_names(gc.get_countries()['NL'], languages=('fr',))
True
An unknown language code contributes nothing rather than raising.
A dictionary keyed by geonameid as a string, holding 34078 cities at the default minimum population of 15000.
>>> gc.get_cities()['2747891']
{
'geonameid': 2747891,
'name': 'Rotterdam',
'latitude': 51.9225,
'longitude': 4.47917,
'countrycode': 'NL',
'population': 868135,
'timezone': 'Europe/Amsterdam',
'admin1code': '11',
'admin2code': '0599',
'featurecode': 'PPL',
'alternatenames': {
'en': ['Rotterdam'],
'es': ['Rotterdam', 'Róterdam'],
'fr': ['Rotterdam'],
'nl': ['Rotterdam']
},
'historicnames': {}
}
alternatenames holds the city's other names grouped by ISO-639 language code, with the preferred name of each language first, and historicnames holds the same for names the source marks as superseded. They follow the rules described under get_admin1_codes(), with one difference: cities keep the languages of their own country plus English, Spanish and French, where divisions keep only their own country's plus English. Untagged names and abbreviations are always kept, and reference codes such as iata or wkdt are excluded. 28753 of the 34078 cities at the default threshold have at least one alternate name, across 208 languages, and 978 have a historic name.
Spanish and French are kept worldwide because their exonyms are what a user types for a city anywhere: "Londres", "Núremberg", "Copenhague". Division names are administrative rather than typed, so they stay scoped to their own country.
Since 5.0 these come from alternateNamesV2.txt rather than the untagged alternatenames column of the cities dumps, which also held GeoNames' ASCII romanisation of every language a place has a name in. Rotterdam's list held 43 entries, 25 of them romanisations such as "Roterdam", "Ratehrdam", "loteleudam" and "rwtrdm", with nothing to say which was which; it now holds only the forms Dutch, English, Spanish and French actually use. Across the default dataset the name count drops from 349573 to 119838, and the bundled city data from 35 MB to 23 MB.
get_city_names() flattens a record the way get_admin1_names() does, name first and deduplicated:
>>> gc.get_city_names(gc.get_cities()['703448'])
['Kyiv', 'Kiev', 'Kijów', 'Kijev', 'Киев', 'Киевом', 'Киеву', 'Київ']
>>> gc.get_city_names(gc.get_cities()['703448'], languages=('uk',))
['Kyiv', 'Київ']
Historic names are separated per language, not per city, and "Kiev" shows why: upstream marks it historic for English but current and preferred for French. So it sits in historicnames['en'] and in alternatenames['fr'] at once, and both searches find the city:
>>> [c['name'] for c in gc.search_cities('Kiev')]
['Kyiv']
>>> [c['name'] for c in gc.search_cities('Kiev', 'historicnames', contains_search=False)]
['Kyiv']
Pass languages to get_city_names() to get one language's view instead.
featurecode is the GeoNames feature code, which distinguishes a capital (PPLC) or an administrative seat (PPLA through PPLA5) from an ordinary populated place (PPL). It is what lets the datasets include capitals below their population threshold, such as Nuuk and Tórshavn. The feature class is always P in these datasets, so it is not stored. You can search on it:
>>> len(gc.search_cities('PPLC', attribute='featurecode', contains_search=False))
241
City names are not unique, so get_cities_by_name() returns a list of records. There is a Rotterdam in both the Netherlands and the US state of New York:
>>> [(c['geonameid'], c['countrycode']) for c in gc.get_cities_by_name('Rotterdam')]
[(2747891, 'NL'), (5134453, 'US')]
Unknown names give an empty list. The first call builds an index of every city name, so looking up many names costs one pass over the dataset instead of one pass per name. get_cities_by_names() returns that whole index, a dictionary mapping each name to its list of records:
>>> len(gc.get_cities_by_names())
32215
search_cities() returns a flat list of city records instead. It searches alternatenames by default, so it matches places whose other names contain the query, here the Rotterdam district of Hoogvliet:
>>> [(c['name'], c['countrycode']) for c in gc.search_cities('Rotterdam')]
[('Rotterdam', 'NL'), ('Hoogvliet', 'NL')]
Pass attribute='name' to search the primary name instead, which finds the US Rotterdam that has no alternate names:
>>> [(c['name'], c['countrycode']) for c in gc.search_cities('Rotterdam', attribute='name')]
[('Rotterdam', 'NL'), ('Rotterdam', 'US')]
First-level administrative divisions (states, provinces, regions), 3865 records keyed by the composite code <countrycode>.<admin1code>, for example US.CA for California or NL.11 for South Holland.
>>> gc.get_admin1_codes()['NL.11']
{
'asciiname': 'Provincie Zuid-Holland',
'geonameid': 2743698,
'name': 'Provincie Zuid-Holland',
'englishname': 'South Holland',
'alternatenames': {
'': ['Zuid-Holland', 'Sudholland', 'Provincie Zuid-Holland', 'South Holland'],
'en': ['South Holland'],
'fy': ['Súd-Hollân'],
'nl': ['Zuid-Holland'],
'abbr': ['zh']
},
'historicnames': {}
}
name and asciiname come from the ADM1 records in the GeoNames allCountries dataset, so they are the current official name and often local-language. The older admin1CodesASCII.txt dataset names divisions after their preferred English alternate name, which upstream lets go stale, e. g. VE.25 is still "Distrito Federal" there years after the rename to "Distrito Capital".
englishname gives that English form back, without the staleness, for the 3381 divisions that have one and as '' for the rest. It is the preferred en alternate name, which is what upstream derives admin1CodesASCII.txt from in the first place, so it reproduces the old names while name stays current:
>>> gc.get_admin1_codes()['VE.25']['name']
'Distrito Capital'
>>> gc.get_admin1_codes()['VE.25']['englishname']
'Distrito Federal'
alternatenames holds the division's other names grouped by ISO-639 language code, with the preferred name of each language first, and historicnames holds the same for names the source marks as superseded. Only the languages of the division's own country (from the Languages column of countryInfo.txt) plus English are kept, along with untagged names and abbreviations. Rows that are reference codes rather than names (link, wkdt, post, iata, icao, faac, unlc, tcid, phon, piny) are excluded.
The language restriction matters because the untagged alternatenames column of allCountries — which this used to be built from — also holds GeoNames' ASCII romanisation of every language it has a name in, and those collide. The Greek for Azerbaijan's Qabala Rayon, Καμπάλα, is also the Greek for Uganda's capital, and romanised to a literal "Kampala" on the division. Reading the language-tagged source keeps Armenian and Russian, which Azerbaijan speaks, in their own scripts instead of as "Gabalayi srjan" and "Gabalinskij rajon".
A division's name plus its alternate names, as one deduplicated list with the name first, the counterpart of get_country_names():
>>> gc.get_admin1_names(gc.get_admin1_codes()['CD.11'])[:4]
['Province du Nord-Kivu', 'Nord-Kivu', 'Sous-Région du Nord-Kivu', 'North Kivu']
Pass languages to take only some of them. Untagged names and abbreviations are always included, whatever languages says, because neither key is a language:
>>> gc.get_admin1_names(gc.get_admin1_codes()['NL.11'], languages=('en',))
['Provincie Zuid-Holland', 'Zuid-Holland', 'Sudholland', 'South Holland', 'zh']
Pass historic=True to append superseded names. They are out by default because a historic name can now belong to somewhere else. Note that GeoNames flags the column sparsely — only 140 of the 3865 divisions have any — so an unflagged name is not evidence that a name is current:
>>> 'Swan River Colony' in gc.get_admin1_names(gc.get_admin1_codes()['AU.08'], historic=True)
True
Second-level administrative divisions (counties, municipalities, districts), 47592 records keyed by <countrycode>.<admin1code>.<admin2code>.
Since 5.0 these are built from the ADM2 rows of allCountries.txt rather than admin2Codes.txt, which carried only a code, a name and an id. The keys and names are unchanged — all 47592 of them — but each record now also carries its code parts and the same language-keyed alternatenames / historicnames / englishname as an admin1 record. 30771 divisions (65%) have at least one alternate name, which is often the only form a reader would write: NG.48.29003 is stored as "Akoko South East" and reported as "Akoko South-East".
>>> gc.get_admin2_codes()['NL.11.0599']
{
'admin1code': '11',
'admin2code': '0599',
'asciiname': 'Rotterdam',
'countrycode': 'NL',
'geonameid': 2747890,
'latitude': 51.9225,
'longitude': 4.47917,
'name': 'Rotterdam',
'englishname': '',
'alternatenames': {...},
'historicnames': {},
}
Note the geonameid here is the municipality of Rotterdam (2747890), which is a different place from the city of Rotterdam (2747891).
Cities store countrycode, admin1code and admin2code separately, so resolving a division means joining them into the composite key. These helpers do that and handle the cases where a city has no code:
>>> city = gc.get_cities()['2747891']
>>> gc.get_admin1_by_city(city)['name']
'Provincie Zuid-Holland'
>>> gc.get_admin2_by_city(city)['name']
'Rotterdam'
Both return None when the city lacks the required codes or the composite key is not in the division dataset, which is why the return value should be checked before subscripting it:
admin1 = gc.get_admin1_by_city(city)
region = admin1['name'] if admin1 else 'unknown'
Building the key by hand works too, but silently produces a partial key such as 'NL.' for cities without an admin1code, so prefer the helpers.
Time zones with their UTC offsets, 418 records keyed by IANA time zone id.
>>> gc.get_timezones()['Europe/Amsterdam']
{
'countrycode': 'NL',
'timezoneid': 'Europe/Amsterdam',
'gmtoffset': 1.0,
'dstoffset': 2.0,
'rawoffset': 1.0
}
rawoffset is the offset excluding daylight saving time. gmtoffset and dstoffset are the offsets in effect on 1 January and 1 July of the year the dataset was published, so they are a snapshot rather than a live value; use a proper time zone library such as zoneinfo if you need the offset at a given moment.
The timezone field of every city record is a key into this dictionary:
>>> city = gc.get_cities()['2747891']
>>> gc.get_timezones()[city['timezone']]['rawoffset']
1.0
The time zones of one country as a list sorted by time zone id. The country code is an ISO alpha-2 code and is matched case insensitively.
>>> [tz['timezoneid'] for tz in gc.get_timezones_by_country('NL')]
['Europe/Amsterdam']
>>> len(gc.get_timezones_by_country('US'))
29
Unknown country codes return an empty list rather than raising:
>>> gc.get_timezones_by_country('ZZ')
[]
A dictionary keyed by the two-letter state code.
>>> gc.get_us_states()['CA']
{'code': 'CA', 'name': 'California', 'fips': '06', 'geonameid': 5332921}
get_us_states_by_names() returns the same records keyed by state name, e. g. gc.get_us_states_by_names()['California'].
A list of 3235 county records, not a dictionary, sourced from the US Census Bureau rather than GeoNames.
>>> gc.get_us_counties()[0]
{'fips': '01001', 'name': 'Autauga County', 'state': 'AL'}
To look counties up, key the list yourself:
counties = {c['fips']: c for c in gc.get_us_counties()}
counties['06037']['name'] # 'Los Angeles County'
The mappers module provides function(s) to map data properties. Currently you can create a mapper that maps country properties, e. g. the name property to the iso3 property, to do so you'd write the following code:
from geonamescache.mappers import country
mapper = country(from_key='name', to_key='iso3')
iso3 = mapper('Spain') # iso3 is assigned ESP
Please write test(s) for any new feature. If you wish to build the data from scratch, run make dl and make json. Note that make dl fetches alternateNamesV2.zip for the country alternate names, which is about 200 MB zipped and 780 MB extracted; ./bin/countries.py streams it once and keeps only the ~250 country records. The bin/ scripts write plain JSON into datasets/, and bin/compress_data.py gzips it into geonamescache/data/ as the last step of make json.