A First Look at the CRU Station List

On Sept 28, 2006, Willis Eschenbach sent an FOI for CRU station data. A year later, after many letters, we still do not have the CRU data as used, but do have a list of stations used, a list which is slightly shorter (4138) than the 4349 stations said to have been used in Brohan et al 2006, their most recent publication. It looks as though they sanitized the list somewhat – in Brohan et al 2006, they said that they removed 55 duplicates. I guess that they’ve identified 156 more duplicates (4 times as many as reported in Brohan et al.) But perhaps there’s another reason. In my opinion, they should have delivered a list of 4349 stations – I’ve asked for this from Phil Jones.

Secondly, the list has not been delivered in working order: the data is supposed to be mostly (“98%”) derived form GHCN, but the identification numbers do not tie in precisely to GHCN numbers. The CRU ID numbers are 6 digits versus 11 digits at GHCN. Many of the CRU numbers tie in to GHCN numbers as follows: the GHCN number is in the form CCCWWWWWDDD where CCC is the country code, WWWWW is the WMO number and DDD is the station – for example, nearby sites (but different) sites can have the same WMO number. GHCN DDD identifiers seldom get out of single digits. CRU identifications of WWWWWD tie to GHCN numbers for 1782 sites. As seen below, many sites can be identified with GHCN sites, but not without a further concordance. CRU says that they have a look-up table, but failed to disclose it. I’ve requested it.

Thirdly, and somewhat unbelievably, the CRU identifications are not unique. In one case, there are 6 stations with an identical ID number. Perhaps there is some still undisclosed list and they’ve delivered it in non-working form for some reason of their own. IF there is no proper list, then I have no idea how they can define a look-up table that functions for non-unique ID numbers. For my own attempt at a concordance, I’ve added a duplicate number for each group of sites with overlapping ID numbers so that the new ID number is unique. (I wasted a considerable effort before I figured out that they had non-unique ID numbers – imagine.)

After doing this, in order to make a concordance, for each CRU site unmatched in a first pass, I then selected GHCN stations that were within 1 degree latitude and within 1 degree longitude and had the first 6 letters identical. If there was only one, I declared a match and assigned the GHCN number. This didn’t match as many sites as could be matched, but reduced the unmatched sites to about 354 sites.

I then made an ASCII tab-separated table in which I wrote down the CRU station plus lat, long and altitude and GHCN stations within a degree plus the same information and manually inspected the stations. I probably could have figured another semi-automatic method of reducing it further, but I also wanted to inspect the matches. In many cases, there was a fairly obvious match with the previous method failing due to multiple candidates or spelling variations. In this way, I added 175 matches, getting up to 3959 matches (a little under 96%), leaving 179 unmatched.

Here are ASCII files listing all CRU stations together with proposed GHCN identifications (ID, name, lat, long shown from GHCN) and the unmatched list. All Unmatched These are ASCII tab-separated and can be opened in Excel or read in R.


Unmatched Stations

The distribution of unmatched stations is really very strange – and I add here, that, for each country specified below, I’ve double checked manually against the GHCN inventory to confirm that there was at least one unmatched CRU station from that country. In total, I identified 29 countries where there were CRU stations that were not present at GHCN, including surprisingly stations from Canada, Australia and even the U.S. Here is a list of countries with at least one station that is unmatched in GHCN:

Argentina: CRU had quite a few stations not at GHCN.
Australia: nearly all match, but two CRU stations didn’t match GHCN – Maryborough and Brisbane Airport. Why these?
Austria – quite a few stations not at GHCN
Bolivia – a couple didn’t match
Brazil – a couple didn’t match
Canada – quite a few stations not at GHCN. I noticed a duplicate GHCN for Parry Sound, which is near Toronto and which occurs in two alter egos in GHCN.
Chile – a couple didn’t match. A couple were called “UNKNOWN” in the CRU list. Perhaps they are connected to the UCAR “Bogus Stations”.
China – quite a few stations not at GHCN
Denmark – a couple didn’t match
Dominica – a couple didn’t match
Germany – one didn’t match
Finland – one (Kuopio) didn’t match
Greenland – possibly a couple didn’t match
Guinea – one didn’t match
Iran – one didn’t match
Ireland – one may not match (Phoenix Park)
Israel – a few don’t match
Italy – a couple don’t match
Kyrgyz republic – one doesn’t match
Netherlands – a couple don’t match
Norway – a few don’t match
Oceania — a few don’t match
Peru – one doesn’t match
Sweden – a few don’t match
Syria – several don’t match
Taiwan – quite a few don’t match
UK – a couple don’t match (Kirkwall, Wick)
USA – about 25 don’t match e.g. Moroni, Lahontan
Russia – a couple may not match

IT is quite weird to see these oddball stations crop at CRU. I’m sure we’ll quickly track down where Moroni and Lahontan and their ilk come from, but it doesn’t seem to be GHCN.

Prior Excuses

With these results in mind, let’s review the history of CRU excuses as to why they should not be required to disclose information under the FOI Act – and it’s taken slightly over a year and many letters and appeals to even get this station list. Their original refusal CRU stated that the data was already located at GHCN as follows:

Datasets named ds564.0 and ds570.0 can be found at The Climate & Global Dynamics Division (CGD) page of the Earth and Sun Systems Laboratory (ESSL) at the National Center for Atmospheric Research (NCAR) site at: http://www.cgd.ucar.edu/cas/tn404/ Between them, these two datasets have the data which the UEA Climate Research Unit (CRU) uses to derive the HadCRUT3 analysis. The latter, NCAR site holds the raw station data (including temperature, but other variables as well). The GHCN would give their set of station data (with adjustments for all the numerous problems). They both have a lot more data than the CRU have (in simple station number counts), but the extra are almost entirely within the USA. We have sent all our data to GHCN, so they do, in fact, possess all our data.

In accordance with S. 17 of the Freedom of Information Act 2000 this letter acts as a Refusal Notice, and the reasons for exemption are as stated below

In response to a further request trying to pin them down, they stated that “more than 98%” of CRU data and the remaining 2% was collected under confidentiality agreements.

Our estimate is that more than 98% of the CRU data are on these sites. The remaining 2% of data that is not in the websites consists of data CRU has collected from National Met Services (NMSs) in many countries of the world. In gaining access to these NMS data, we have signed agreements with many NMSs not to pass on the raw station data, but the NMSs concerned are happy for us to use the data in our gridding, and these station data are included in our gridded products, which are available from the CRU web site. These NMS-supplied data may only form a very small percentage of the database, but we have to respect their wishes and therefore this information would be exempt from disclosure under FOIA pursuant to s.41. The World Meteorological Organization has a list of all NMSs.

Obviously, none of this justified not providing a list of stations, but that has taken another 6 months. In connection with the supposed confidentiality agreements, as reported previously, Doug Keenan asked for the countries with which there were confidentiality agreements that restricted access and was told:

Dear Doug,
I have done some searching in files – all from the period 1990-1998. This is the time when we were in contact with a number of NMSs. We have also got datasets from fellow scientists and other institutes around the world. All supplied data (eventually and sometimes at cost), but we were asked not to pass on the raw data to third parties, but we could use the data to develop products (our datasets) and use the data in scientific papers. It is likely that some of the NMSs and Institutes have changed their policies now – and that the people we were corresponding with (all by regular mail or fax) are no longer there or are in different sections. The lists below don’t refer to all the stations within these countries, nor to all periods, but to some of the data for some of the time.
The NMSs
Germany, Bahrain, Oman, Algeria, Japan, Slovakia and Syria

Scientists/Institutes (data for these countries)
Mali, India, Pakistan, Poland, Indonesia, Democratic Republic of the Congo (was Zaire), Sudan and some Caribbean Islands.

These are the only ones I can find evidence for. I’m sure there were a few others during the 1980s, but we have moved buildings twice since 1980.

Not sure how you will use this data.
Phil Jones

Above I summarized the countries for which there are stations that are not matched at GHCN. Remarkably, these include virtually none of the countries where Jones said that they had received data subject to confidentiality agreements – so that the confidentiality agreement excuse cannot apply for any of these countries. And for each of the countries for which Jones said that there was a confidentiality agreement (Bahrain, Oman, Algeria, Japan, Slovakia, Mali, India, Pakistan, Poland, Indonesia, Zaire and Sudan), I was able to cross-identify all CRU stations with GHCN identifications so that the confidentiality excuse didn’t affect anything.

At this point, the only unmatched stations which would appear to be covered by a reported “confidentiality agreement” are about 6 stations in Syria (about half at GHCN) and one German station (Wahnsdorff). Otherwise there is no valid excuse for not disclosing this station data. Of course it is possible that Jones has confidentiality agreements with Canada and the Australia, but was embarrassed to report them and thus omitted them in the above list. We’ll see.

It is disappointing that the pretexts for not providing the data previously have turned out to be untrue. However, it should be possible to now develop a reasonable concordance for the CRU stations to GHCN where applicable and to identify provenances for the oddball stations to make a concordance up to a very small number of stations – at which analysis can begin.

CRU Reveals Station Identities

Willis Eschenbach received a message today that the CRU list of stations used is online at http://www.cru.uea.ac.uk/cru/data/landstations/ . The webpage says:

The file gives the locations and names of the 4138 stations used at some time (i.e. in the gridding that is used to produce CRUTEM3) during the period from 1851 to 2006. All these stations have adequate 30 year averages for 1961-90 as defined in Brohan et al. (2006). The 4138 total is lower than the 4349 value given as the starting point for Brohan et al. (2006) and used in the latest IPCC Report. A small number of stations have been removed during Brohan et al. (2006) because of the presence of duplicate data and insufficient coverage for the period 1961-90.

They say:

The numbers we use are listed in numerical order up to station number 988360. Up to this point, the numbers ending in zeroes are generally the WMO number (*10) in use for that station in the mid-1980s. Numbers not ending in a zero have generally been assigned by CRU or may have originated from other sources. Stations that are listed after number 988360 are stations for which CRU has assigned numbers, mostly beginning with 72 (so using spare country numbers not officially used by WMO) to 75 (corresponding to stations in the United States). Some WMO IDs have been updated in the 2000s.

It looks like the sixth digit is the GHCN identification. This seems to be the case with Marysville which I checked. They note:

The station temperature data are updated each month, together with some back data for the last couple of years. As the WMO IDs have not all been updated, we have a look-up Table which associates some current WMO station numbers with the earlier values we are using. Updates come from two principal sources [CLIMAT messages exchanged between National Meteorological Services (NMSs) and from the publication Monthly Climatic Data for the World]. Additional updates in near-real time (either monthly or annually) come directly from Australia, Canada, New Zealand, Austria, the Nordic countries and a few others.

The look-up table doesn’t seem to be online.

Some aspects of HadCRU3 are easier to implement than GISS – earlier this year, I was able to track a gridcell series to the Barabinsk station data. With Hansen, you can never track anything. There’s a case for putting some of the Hansen mysteries on hiatus for a while and digging into CRU. The effort in understanding the individual stations is not lost since we’re mostly dealing with GHCN and MCDW data.

Dongge Cave

Dongge Cave is a very long speleothem in southeast China, which is held to provide evidence on changes in the Asian monsoon.

It was most recently considered in Bao Yang et al (QSR 2007) in juxtaposition with the Dasuopu ice core, Oman speleothem S3 and the RC2730 Arabian Sea core showing G bulloides percentage. (This latter core was an important contributor to the Moberg reconstruction and has been discussed on several occasions – G bulloides are evidence for cold water and are interpreted as evidence of increased upwelling of cold water, which in turn is considered to be evidence of wind speed, and thus monsoon activity.)

Kim Cobb has also just reported a new speleothem in Borneo, somewhat the same part of the world, which I’ll try to look at soon. I’ve not previously looked at the Dongge Cave speleothem. Its results have been properly archived at WDCP, much facilitating examination. Continue reading →

A Second Look at USHCN Classification #2

Continuation

Unthreaded #21

No discussion of CO2 measurements, thermodynamics, theory of radiation, etc. please – other than to identify interesting references – and something more than the title is usually helpful. How hard can that be? If anyone can identify a clear exposition of how 2xCO2 leads to 2.5 deg C, please do so. (I’m not taking any position on the matter, I’m just trying to identify the best possible exposition. )

Russian Bias

There has been much discussion on this site regarding the methodology employed by Dr. Hansen’s “Step 1”, also known as the “bias method” in HL87. Readers unfamiliar with the topic may want to read through the material here, here, and here to gain a better understanding of the method and issues raised with employing it.

What was clear since we first unraveled the process was that it was destined to corrupt the combined station data and, as it’s name implies, add a bias to that combined data. What was not clear was what the net effect would be to the regional and worldwide temperature record.

Thanks to Steve’s help, I recently completed an initial look into the effect the bias method has on the temperature record.

Continue reading →

YTD Hurricane Activity

Some of you have been noticing a tendency for almost any gust of wind in the Atlantic to now become a named storm. Given this tendency, more relevant metrics are obviously the number of hurricane-days (and the closely related ACE index) and the number of storm-days.

I’ve scraped the data and done the YTD calculations, comparing these to the corresponding values to the end of September in previous years (I’ll replace this graphic in a few days when Sept 2007 is completed, but I don’t expect much change.) Continue reading →

Houston, We've Found Wellington NZ

As noted before, climateaudit readers have helped UCAR find the lost civilization of Chile and today, we are happy to report that we have helped NASA find the lost city of Wellington NZ.

NASA’s records for Wellington NZ were mysteriously interrupted in 1988 – an interruption so severe that we assumed that Wellington NZ must have been destroyed by Scythians. We are happy to report that Wellington NZ is still in existence.

This is not the only good news. We are also happy to announce that there is still a functioning meteorological service. Not only that, but can announce contact with the indigenous representatives.

Although NASA (and NOAA) appear to have lost contact, an indigenous NZ climate scientist familiar with the lost records has contacted climateaudit. I have passed this exciting news on to NASA and urged them to restore contact with their lost cousins in NZ. Continue reading →

TOBS

It’s true that TOB is pretty far from the topic of this thread, so perhaps our host could start a new one just on TOB, ideally copying into it the pertinent posts from this thread? I may have missed a few, but a good start would be #305, 376, 400, 402, 403, 413, 418, 419, 420, 424, 455, 458, 460, 462, 464, 468, 484, 488, and 493. )

Hugues Goosse and the Unresponsiveness of Juckes

On Sep 21, 2007, Hugues Goosse, the Climate of the Past editor responsible for the Juckes article, published a statement saying that a revised version of the Juckes et al article had been submitted to “conventional” refereeing and accepted on Sept 21, 2007. He said:

On the other hand, the authors disagree with one reviewer on some points for which no clear consensus could be gained from published literature. The arguments of the authors appear reasonable from our present knowledge of the field and are presented in a balanced way. As a consequence, I decided to accept the paper for publication in Climate of the Past.

I presume that I was the “one reviewer”, although Willis Eschenbach and Mark Rostron also submitted critical reviews. Under CP policies, authors are supposed to respond to review comments. I’ve collated my review comments together with Juckes’ replies. It is remarkable how insolent and unresponsive Juckes’ comments are.

In virtually every case, I’ve provided a detailed and analytical comment and Juckes virtually never makes a direct and straightforward reply, rebutting the comment in straightforward terms. See what you think. Continue reading →