Steig’s Trick

Please note:  The author of this post is Ryan O . . . not Steve.  As people have lately displayed a tendency to attribute what I write to Steve, I figured the disclaimer was appropriate.

Steve: Feb 9, 2011 – some of Ryan’s language, including the original title, breached blog policies and has been edited accordingly.
***

Some of you may have noticed that Eric Steig has a new post on our paper at RealClimate.  In the past when I have wished to challenge Eric on something, I generally have responded at RealClimate.  In this case, a more detailed response is required, and a simple post at RC would be insufficient.  Based on the content, it would not have made it past moderation anyway.

Lest the following be entirely one-sided, I should note that most of my experiences with Eric in the past have been positive.  He was professional and helpful when I was asking questions about how exactly his reconstruction was performed and how his verification statistics were obtained.  My communication with him following acceptance of our paper was likewise friendly.  While some of the public comments he has made about our paper have fallen far short of being glowing recommendations, Eric has every right to argue his point of view and I do not begrudge his doing so.  I should also note that over the past week I was contacted by an editor from National Geographic, who mentioned in passing that he was referred to me by Eric.  This was quite gracious of Eric, and I honestly appreciated the gesture.

However, once Eric puts on his RealClimate hat, his demeanor is something else entirely.  Again, he has every right to blog about why he feels our paper is something other than how we have characterized it (just as we have every right to disagree).  However, what he does not have the right to do is to defend his point of view by [snip] misrepresenting facts.

In other words, in his latest post, Eric crossed the line.

Let us examine how (with the best, of course, saved for last).

***

The first salient point is that Eric still doesn’t get it.  The whole purpose of our paper was to demonstrate that if you properly use the data that S09 used, then the answer changes in a significant fashion.  This is different than claiming that this particular method (whereby satellite data and ground station data are used together in RegEM) provides a more accurate representation of the [unknown] truth than other methods.  We have not (and will not) make such a claim.  The only claim we make is – given the data and regression method used by S09 – that the answer is different when the method by which the data are combined is properly employed.  Period.

The question about whether the proper use of the AVHRR and station data sets yield an accurate representation of the temperature history of Antarctica is an entirely separate topic.  To be sure, it is an important one, and it is a legitimate course of scientific inquiry for Eric to argue that our West Antarctic results are incorrect based on independent analyses.  What is entirely, wholly, and completely not legitimate is to use those same arguments to defend the method of his paper, as the former makes no statement on the latter.

Unfortunately, Eric does not seem to understand.  He wishes to continue comparing our results to other methods and data sets (such as NECP, ERA-40, Monaghan’s kriging method, and boreholes).  We did not use those sets or methods, and we make no comment on whether analyses conducted using those sets and methods are more likely to give better results.  Yet Eric insists on using such comparisons to cast doubt on our methodological criticisms of the S09 method.

While such comparisons are, indeed, important for determining what might be the true temperature history of Antarctica, they have absolutely nothing to do with the criticisms advanced in our paper.  Zero.  Zilch.  Nada.  Note how Eric has refrained from talking about those criticisms directly.  I can only assume that this is because he has little to say, as the criticisms are spot-on.

Instead, what Eric would prefer to do is look at other products and say, “See!  Our West Antarctic trends at Byrd Station are closer than O’Donnell’s!  We were right!”  While it may be a true statement that the S09 results at Byrd Station prove to be more accurate as better and better analyses are performed, if so, it was sheer luck (as I will demonstrate, yet again).  The S09 analysis does not have the necessary geographic resolution nor the proper calibration method to independently demonstrate this accuracy.

I could write a chapter in the Farmer’s Almanac explaining how the global temperature will drop by 0.5 degrees by 2020 and base my analysis on the alignment of the planets and the decline in popularity of the name “Al”.  If the global temperature drops by 0.5 degrees by 2020, does that validate my method – and, by extension, invalidate the criticisms against my method?  Eric, apparently, would like to think so.

If he wishes to argue that our results are incorrect, that’s fine.  To be quite honest, I would hope that he would do exactly that if he has independent evidence to support his views (and he does, indeed, have some).  But if he wishes to defend his method, then it is time for him to begin advancing mathematically correct arguments why his method was better (or why our criticisms were not accurate).  Otherwise, it is time for Eric to stop playing the carnival prognosticator’s game of using the end result to imply that an inappropriate use of information was somehow “right” because – by chance – the answer was near to the mark.

***

The second salient point relates to the evidence Eric presents that our reconstruction is less accurate.  When it comes to differences between the reconstruction and ground data, Eric focuses primarily on Byrd station.  While his discussion seems reasonable at first glance, it is quite misleading.  Let us examine Eric’s comments on Byrd in detail.

Eric first presents a plot where he displays a trendline of 0.38 +/- 0.2 Deg C / decade (Raw data) for 1957 – 2006.  He claims that this is the ground data (in annual anomalies) for Byrd station.  While there is not much untrue about this statement, there is certainly a [material] [snip] omission.  To see this [material omission], we only need look at the raw data from Byrd over this period:

Pay close attention to the post-2000 timeframe.  Notice how the winter months are absent?  Now what do we suppose might happen if we fit a trend line to this data?  One might go so far as to say that the conclusion is foregone.

So . . . would Eric Steig really do this?  [snip]

Trend check on the above plot:  0.38 Deg C / decade.

Hm.

By the way, the trend uncertainty when the trend is calculated this way is +/- 0.32, and, if one corrects for the serial correlation in the residuals, it jumps to +/- 0.86.  But since neither of those tell the right story, I suppose the best option is to simply copy over the +/- 0.2 from the Monaghan reconstruction trend, or the uncertainty from the properly calculated trend.

When calculated properly, the 50-year Byrd trend is 0.25 +/- 0.2 (corrected for serial correlation).  This is still considerably higher than the Byrd location in our reconstruction, and is very close to the trend in the S09 reconstruction.  However, we’ve yet to address the fact that the pre-1980 data comes from an entirely different sensor than the post-1980 data.

Eric notes this fact, and says:

Note that caution is in order in simply splicing these together, because sensor calibration issues could means that the 1°C difference is an overestimate (or an underestimate).

He then proceeds to splice them together anyway.  His justification is that there is about a 1oC difference in the raw temperatures, and then goes on to state that there is independent evidence from a talk given at the AGU conference about a borehole measurement from the West Antarctic Ice Sheet Divide.  This would be quite interesting, except that the first half of the statement is completely untrue.

If you look at the raw temperatures (which was how he computed his trend, so one might assume that he would compare raw temperatures here as well), the manned Byrd station shows a mean of -27.249 Deg C.  The AWS station shows a mean of -27.149 . . . or a 0.1 Deg C difference.  Perhaps he missed a decimal point.

However, since computing the trend using the raw data is unacceptable due to an uneven distribution of months for which data is present, computing the difference in temperature using the raw data is likewise unacceptable.  The difference in temperature should be computed using anomalies to remove the annual cycle.  If the calculation is done this (the proper) way, the manned Byrd station shows an average anomaly of -0.28 Deg C and the AWS station shows an average of 0.27 Deg C.  This yields a difference of 0.55 Deg C . . . which is still not 1 Deg C.

Maybe he’s rounding up?

Seriously, Eric . . . are we playing horseshoes?

Of course, of this 0.55 Deg C difference, fully one-third is due to a single year (1980), which occurred 30 years ago:

Without 1980 – which occurs at the very beginning of the AWS record – the difference between the manned station anomalies and the AWS anomalies is 0.37 Deg C.

Furthermore, even if there were a 1 degree difference in the manned station and AWS values, this still doesn’t tell the story Eric wants it to.

The original Byrd station was located at 119 deg 24’ 14” W.  The AWS station is located at 119 deg 32’ W.  Seems like almost the same spot, right?  The difference is only about 2.5 km.  This is the same distance as that between McMurdo (elev. 32m) and Scott Base (elev. 20m).  So if one can willy-nilly splice the Byrd station together, one would expect that the same could be done for McMurdo and Scott Base.  So let’s look at the mean temperatures and trends for those two stations (both of which have nearly complete records).

McMurdo:  -16.899 Deg. C (mean)

Scott Base:  -19.850 Deg. C (mean)

That’s a 3 degree difference for stations at a similar elevation and a linear separation of 2.5 km . . . just like Byrd manned and Byrd AWS.

So what would the trend be if we spliced the first half of Scott Base with the second half of McMurdo?

1.05 +/- 0.19 Deg C / decade.

OMG . . . it is SO much worse than we thought!

Microclimate matters.  Sensor differences matter.  The fact that AWS stations are likely to show a warming bias compared to manned stations (as the distance between the sensor and the snow surface tends to decrease over time, and Antarctica shows a strong temperature gradient between the nominal 3m sensor height and the snow surface) matters.  All of these are ignored by Eric, and he should know better.

Eric goes on to state that this meant we somehow used less available station data than he did:

On top of that, O’Donnell et al. do not appear to have used all of the information available from the weather stations. Byrd is actually composed of two different records, the occupied Byrd Station, which stops in 1980, and the Byrd AWS station which has episodically recorded temperatures at Byrd since then. O’Donnell et al. treat these as two independent data sets, and because their calculations (like ours) remove the mean of each record, O’Donnell et al. have removed information that might be rather important. namely, that the average temperatures in the AWS record (post 1980) are warmer — by about 1 Deg C — than the pre-1980 manned weather station record.

In reality, the situation is quite the opposite of what Eric implies.  We have no a priori knowledge on how the two Byrd stations should be combined.  We used the relationships between the two Byrd stations and the remainder of the Antarctic stations (with Scott Base and McMurdo – which show strong trends of 0.21 and 0.25 Deg C / decade – dominating the regression coefficients) to determine how far the two stations should be offset.  By simply combining the two stations without considering how they relate to any other stations, it was Eric who threw this information away.

With this being said, combining the two station records without regard to how they relate to other stations does change our results.  So if Eric could somehow justify doing so, our West Antarctic trend would increase from 0.10 Deg C / decade to 0.16 Deg C / decade, and the area of statistically significant trends would grow to cover the WAIS divide, yielding statistically significant warming over 56% (instead of 33%) of West Antarctica.  However, to do so, Eric must justify why it is okay to allow RegEM to determine offsets for every other infilled point in Antarctica except Byrd, and furthermore must propose and justify a specific value for the offset.  If RegEM cannot properly combine the Byrd stations via infilling missing values, then what confidence can we have that it can properly infill anything else?  And if we have no confidence in RegEM’s ability to infill, then the entire S09 reconstruction – and, by extension, ours – are nothing more than mathematical artifacts.

However, to guard against this possibility (unlike S09), we used an alternative method to determine offsets as a check against RegEM (credit Jeff Id for this idea and the implementation).  Rather than doing any infilling, we determined how far to offset non-overlapping stations by comparing mean temperatures between stations that were physically close, and using these relationships to provide the offsets.  This method yielded patterns of temperature change that were nearly identical to the RegEM-infilled reconstructions, with a resulting West Antarctic trend of 0.12 Deg C / decade.

Lastly, Eric implies that his use of Byrd as a single station somehow makes his method more accurate.  This is hardly true.  Whether you use Byrd as a single station or two separate stations, the S09 answer changes by a mere 0.01 Deg C / decade in West Antarctica and 0.005 Deg C / decade at the Byrd location.  The characteristic of being entirely impervious to changes in the “most critical” weather station data is a rather odd result for a method that is supposed to better utilize the station data.

Interestingly, if you pre-combine the Byrd data like Eric does and perform the reconstruction exactly like S09, the resulting infilled ground station trend at Byrd is 0.13 Deg C / Decade (fairly close to our gridded result).  The S09 gridded result, however, is 0.25 Deg C / Decade – or almost double the ground station trend from their own RegEM infilling, and closer to our gridded result than to theirs.

(Weird.  Didn’t Eric say their reconstruction better captures the ground station data from Byrd?  Hm.)

Stranger yet, if you add a 0.1 Deg C / decade trend to the Peninsula stations, the S09 West Antarctic trend increases from 0.20 to 0.25 – with most of the increase occurring 2,500 km away from the Peninsula on the Ross Ice Shelf – the East Antarctic trend increases from 0.10 to 0.13 . . . but the Peninsula trend only increases from 0.13 to 0.15.  So changes in trends at the Peninsula stations result in bigger changes in West and East Antarctica than in the Peninsula!  Nor is this an artifact of retaining the satellite data (sorry, Eric, but I’m going to nip that potential arm-flailing argument in the bud).  Using the modeled PCs instead of the raw PCs in the satellite era, the West trend goes from 0.16 to 0.22, East goes from 0.08 to 0.12, and the Peninsula only goes from 0.11 to 0.14.

(Weird.  Didn’t Eric say their reconstruction better captures the ground station data from Byrd?  Hm.)

And (nope, still not done with this game) EVEN STRANGER YET, if you add a whopping 0.5 Deg C / decade trend to Byrd (or five times what we added to the Peninsula), the S09 West Antarctic trend changes by . . . well . . . a mere 0.02 Deg C / decade, and the gridded trend at Byrd station rises massively from 0.25 to . . . well . . . 0.29.  The Byrd station trend used to produce this result, however, clocks in at a rather respectable 0.75 Deg C / decade.  Again, this is not an artifact of retaining the satellite data.  If you use the modeled PCs, the West trend increases by 0.03 and the gridded trend at Byrd increases to only 0.26.

(Weird.  Didn’t Eric say their reconstruction better captures the ground station data from Byrd?  Hm.)

Now, what happens to our reconstruction if you add a 0.1 Deg C / decade trend to the Peninsula stations?  Our East Antarctic trends go from 0.02 to . . . 0.02.  Our West Antarctic trends go from 0.10 to 0.16, with almost all of the increase in Ellsworth Land (adjacent to the Peninsula).  And the Peninsula trend goes from 0.35 to 0.45 . . . or the same 0.1 Deg C / decade we added to the Peninsula stations.

And what happens if you add a 0.5 Deg C / decade trend to Byrd?  Why, the West Antarctic trend increases 160% from 0.10 to 0.26 Deg C / decade, and the gridded trend at Byrd Station increases to 0.59 Deg C / decade . . . with the East Antarctic trends increasing by a mere 0.01  and the Peninsula trends increasing by 0.03.

(Weird.  Didn’t Eric say their reconstruction better captures the ground station data from Byrd?  Hm.)

***  You see, Eric, the nice thing about getting the method right is that if the data changes – or more data becomes available (like, say, a better way to offset Byrd station than using the relationships to other stations), then the answer will change in response.  So if someone uses our method with better data, they will get a better answer.  If someone uses your method with better data, well, they will get the same answer . . . or they will get garbage.  This is why I find this comment by you to be particularly ironic:  ***

At some point, yes. It’s not very inspiring work, since the answer doesn’t change [indeed; your method is peculiarly robust to changes in the data it supposedly represents], but i suppose it has to get done. I had hoped O’Donnell et al. would simply get it right, and we’d be done with the ‘debate’, but unfortunately not.—eric

(emphasis and bracketed text added by me)

Eric’s claims that his reconstruction better captures the information from the station data are wholly and demonstrably false.  With about 30 minutes of effort he could have proven this to himself . . . not only for his reconstruction, but also for ours.  This is likely to be less time than it took him to write that post.  You would think that if he felt strongly enough about something to make a public critique that he would have taken the time to verify whether any of his suppositions were correct.  This is apparently not the case.  On that note, I found this comment by Eric to be particularly infuriating:

If you can get their code to work properly, let me know. It’s not exactly user friendly, as it is all in one file, and it takes some work to separate the modules.

Here’s how you do it, Eric:

  1. Go to CRAN (http://cran.r-project.org/) and download the latest version of R.
  2. Download our code here:  http://www.climateaudit.info/data/odonnell
  3. Open up R.
  4. Open up our code in Notepad.
  5. Put your cursor at the very top of our code.
  6. Go all the way to the end of the code and SHIFT-CLICK.
  7. Press CTRL-C.
  8. Go to R.
  9. Press CTRL-V.
  10. Wait about 17 minutes for the reconstructions to compute.

Easy-peasy.  You didn’t even try.

I am sick of arm-waving arguments, unsubstantiated claims, and uncalled-for snark.  Did you think I wouldn’t check?  I would have thought you would have learned quite the opposite from your experience reviewing our paper.

Oops.

Did I let something slip?

***

I mentioned at the beginning that I was planning to save the best for last.

I have known that Eric was, indeed, Reviewer A since early December.  I knew this because I asked him.  When I asked, I promised that I would keep the information in confidence, as I was merely curious if my guess that I had originally posted on tAV had been correct.

Throughout all of the questioning on Climate Audit, tAV, and Andy Revkin’s blog, I kept my mouth shut.  When Dr. Thomas Crowley became interested in this, I kept my mouth shut.  When Eric asked for a copy of our paper (which, of course, he already had) I kept my mouth shut.  I had every intention of keeping my promise . . . and were it not for Eric’s latest post on RC, I would have continued to keep my mouth shut.

However, when someone makes a suggestion during review that we take and then later attempts to use that very same suggestion to disparage our paper, my obligation to keep my mouth shut ends.

(Note to Eric:  unsubstantiated arm-waving may frustrate me, but [snip] is intolerable.)

Part of Eric’s post is spent on the choice to use individual ridge regression (iRidge) instead of TTLS for our main results.  He makes the following comment:

Second, in their main reconstruction, O’Donnell et al. choose to use a routine from Tapio Schneider’s ‘RegEM’ code known as ‘iridge’ (individual ridge regression). This implementation of RegEM has the advantage of having a built-in cross validation function, which is supposed to provide a datapoint-by-datapoint optimization of the truncation parameters used in the least-squares calibrations. Yet at least two independent groups who have tested the performance of RegEM with iridge have found that it is prone to the underestimation of trends, given sparse and noisy data (e.g. Mann et al, 2007a, Mann et al., 2007b, Smerdon and Kaplan, 2007) and this is precisely why more recent work has favored the use of TTLS, rather than iridge, as the regularization method in RegEM in such situations. It is not surprising that O’Donnell et al (2010), by using iridge, do indeed appear to have dramatically underestimated long-term trends—the Byrd comparison leaves no other possible conclusion.

The first – and by far the biggest – problem that I have with this is that our original submission relied on TTLS.  Eric questioned the choice of the truncation parameter, and we presented the work Nic and Jeff had done (using ridge regression, direct RLS with no infilling, and the nearest-station reconstructions) that all gave nearly identical results.

What was Eric’s recommendation during review?

My recommendation is that the editor insist that results showing the ‘mostly [sic] likely’  West Antarctic trends be shown in place of Figure 3.  [the ‘most likely’ results were the ridge regression results] While the written text does acknowledge that the rate of warming in West Antarctica is probably greater than shown, it is the figures that provide the main visual ‘take home message’ that most readers will come away with. I am not suggesting here that kgnd = 5 will necessarily provide the best estimate, as I had thought was implied in the earlier version of the text. Perhaps, as the authors suggest, kgnd should not be used at all, but the results from the ‘iridge’ infilling should be used instead. . . . I recognize that these results are relatively new – since they evidently result from suggestions made in my previous review [uh, no, not really, bud . . . we’d done those months previously . . . but thanks for the vanity check] – but this is not a compelling reason to leave this ‘future work’.

(emphasis and bracketed comments added by me)

And after we replaced the TTLS versions with the iRidge versions (which were virtually identical to the TTLS ones), what was Eric’s response?

The use of the ‘iridge’ procedure makes sense to me, and I suspect it really does give the best results. But O’Donnell et al. do not address the issue with this procedure raised by Mann et al., 2008, which Steig et al. cite as being the reason for using ttls in the regem algorithm. The reason given in Mann et al., is not computational efficiency — as O’Donnell et al state — but rather a bias that results when extrapolating (‘reconstruction’) rather than infilling is done. Mann et al. are very clear that better results are obtained when the data set is first reduced by taking the first M eigenvalues. O’Donnell et al. simply ignore this earlier work. At least a couple of sentences justifying that would seem appropriate.

(emphasis added by me)

So Eric recommends that we replace our TTLS results with the ridge regression ones (which required a major rewrite of both the paper and the SI) and then agrees with us that the iRidge results are likely to be better . . . and promptly attempts to turn his own recommendation against us.

There are not enough vulgar words in the English language to properly articulate my disgust [snip].

The second infuriating aspect of this comment is that he tries to again misrepresent the Mann article to support his claim when he already knew [or ought to have known] otherwise. [snip] In the response to the Third Review, I stated:

We have two topics to discuss here.  First, reducing the data set (in this case, the AVHRR data) to the first M eigenvalues is irrelevant insofar as the choice of infilling algorithm is concerned.  One could just as easily infill the missing portion of the selected PCs using ridge regression as TTLS, though some modifications would need to be made to extract modeled estimates for ridge.  Since S09 did not use modeled estimates anyway, this is certainly not a distinguishing characteristic.

The proper reference for this is Mann et al. (2007), not (2008).  This may seem trivial, but it is important to note that the procedure in the 2008 paper specifically mentions that dimensionality reduction was not performed for the predictors, and states that dimensionality reduction was performed in past studies to guard against collinearity, not – as the reviewer states – out of any claim of improved performance in the absence of collinear predictors.  Of the two algorithms – TTLS and ridge – only ridge regression incorporates an automatic check to ensure against collinearity of predictors.  TTLS relies on the operator to select an appropriate truncation parameter.  Therefore, this would suggest a reason to prefer ridge over TTLS, not the other way around, contrary to the implications of both the reviewer and Mann et al. (2008).

The second topic concerns the bias.  The bias issue (which is also mentioned in the Mann et al. 2007 JGR paper, not the 2008 PNAS paper) is attributed to a personal communication from Dr. Lee (2006) and is not elaborated beyond mentioning that it relates to the standardization method of Mann et al. (2005).  Smerdon and Kaplan (2007) showed that the standardization bias between Rutherford et al. (2005) and Mann et al. (2005) results from sensitivity due to use of precalibration data during standardization.  This is only a concern for pseudoproxy studies or test data studies, as precalibration data is not available in practice (and is certainly unavailable with respect to our reconstruction and S09).

In practice, the standardization sensitivity cannot be a reason for choosing ridge over TTLS unless one has access to the very data one is trying to reconstruct.  This is a separate issue from whether TTLS is more accurate than ridge, which is what the reviewer seems to be implying by the term “bias” – perhaps meaning that the ridge estimator is not a variance-unbiased estimator.  While true, the TTLS estimator is not variance-unbiased either, so this interpretation does not provide a reason for selecting TTLS over ridge.  It should be clear that Mann et al. (2007) was referring to the standardization bias – which, as we have pointed out, depends on precalibration data being available, and is not an indicator of which method is more accurate.

More to [what we believe to be] the reviewer’s point, though Mann et al. (2005) did show  in the Supporting Information where TTLS demonstrated improved performance compared to ridge, this was by example only, and cannot therefore be considered a general result.  By contrast, Christiansen et al. (2009) demonstrated worse performance for TTLS in pseudoproxy studies when stochasticity is considered – confirming that the Mann et al. (2005) result is unlikely to be a general one.  Indeed, our own study shows ridge to outperform TTLS (and to significantly outperform the S09 implementation of TTLS), providing additional confirmation that any general claims of increased TTLS accuracy over ridge is rather suspect.

We therefore chose to mention the only consideration that actually applies in this case, which is computational efficiency.  While the other considerations mentioned in Mann et al. (2007) are certainly interesting, discussing them is extratopical and would require much more space than a single article would allow – certainly more than a few sentences.

Note some curious changes from Eric’s review comment and his RC post.  In his review comment, he refers to Mann 2008.  I correct him, and let him know that the proper reference is Mann 2007.  He also makes no mention of Smerdon’s paper.  I do.  I also took the time to explain, in excruciating detail, that the “bias” referred to in both papers is standardization bias, not variance bias in the predicted values.

So what does Eric do?  Why, he changes the references to the ones I provided (notably, excluding the Christiansen paper) and proceeds to misrepresent them in exactly the same fashion that he tried during the review process!  [SM Update Feb 9- Steig stated by email today that he did not see the Response to Reviewer A’s Third Review; the amendment of the incorrect reference in the Third Review to the correct references provided in the Response to the Third Review was apparently a coincidence.]

And by the way, in case anyone (including Eric) is wondering if I am the one who is misrepresenting, fear not.  Nic and I contacted Jason Smerdon by email to ensure our description was accurate.

But the B.S. piles even deeper.  Eric implies that the reason the Byrd trends are lower is due to variance loss associated with iRidge.  He apparently did not bother to check that his reconstruction shows a 16.5% variance loss (on average) in the pre-satellite era when compared to ours. The reason for choosing the pre-satellite era is that the satellite era in S09 is entirely AVHRR data, and is thus not dependent on the regression method.  We also pointed this out during the review . . . specifically with respect to the Byrd station data. Variance loss due to regularization bias has absolutely NOTHING to do with the lower West Antarctic trends in our reconstruction . . . and [snip].

This knowledge, of course, does not seem to stop him from implying the opposite.

Then Eric moves on to the TTLS reconstructions from the SI, grabs the kgnd = 6 reconstruction, and says, “See?  Overfitting!” without, of course, providing any evidence that this is the case.  He goes on to surmise that the reason for the overfitting is that our cross-validation procedure selected the improper number of modes to retain – yet again without providing any evidence that this is the case (other than it better matches his reconstruction).

So if Eric is right, then using kgnd = 6 should better capture the Byrd trends than kgnd = 7, right?  Let’s see if that happens, shall we?

If we perform our same test as before (combining the two Byrd stations and adding a 0.5 Deg C trend, so an initial Byrd trend of 0.75 Deg C / decade), we get:

Infilled trend (kgnd = 6):  0.45 Deg C / decade

Infilled trend (kgnd = 7):  0.52 Deg C / decade

Weird.  It looks as if the kgnd = 7 option better captures the Byrd trend . . . didn’t Eric say the opposite?  Hm.

These translate into reconstruction trends at the Byrd location of 0.42 and 0.45 Deg C / decade, respectively (you can try other, more reasonable trends if you want . . . it doesn’t matter).  I also note that the TTLS reconstructions do a poorer job of capturing the Byrd ground station trend than the iRidge reconstructions, which is the opposite behavior suggested by Eric (and this was noted during the review process as well).

Perhaps Eric meant that we overfit the “data rich” area of the Peninsula?  Fear not, dear Reader, we also have a test for that!  Let’s add our 0.1 Deg C / decade trend to the Peninsula stations, shall we, and see what results:

Recon trend increase (kgnd = 6):  Peninsula +0.08, West +0.02, East +0.02

Recon trend increase (kgnd = 7):  Peninsula + 0.11, West +0.03, East +0.01

Weird.  It looks as if the kgnd = 7 option better captures the Peninsula trend with a similar effect on the East or West trends . . . didn’t Eric say the opposite?  Hm.

By the way, Eric also fails to note that the kgnd = 5 and 6 Peninsula trends, when compared to the corresponding station trends, are outside the 95% CIs for the stations.  I guess that’s okay, though, since the only station that really matters in all of Antarctica is Byrd (even though his own reconstruction is entirely immune to Byrd).

As far as the other misrepresentations go in his post, I’m done with the games.  These were all brought up by Eric during the review.  Rather than go into detail here, I will shortly make all of the versions of our paper, the reviews, and the responses available at http://www.climateaudit.info/data/odonnell.

***

My final comment is that this is not the first time.

At the end of his post, Eric suggests that the interested Reader see his post “On Overfitting”.  I suggest the interested Reader do exactly that.  In fact, I suggest the interested Reader spend a good deal of time on the “On Overfitting” post to fully absorb what Eric was saying about PC retention.  Following this, I suggest that the interested Reader examine my posts in that thread.

Once this is completed, the interested Reader may find Review A rather . . . well . . . interesting when the Reader comes to the part where Eric talks about PC retention.

Fool me once, shame on you.  But twice isn’t going to happen, bud.

Sci Tech Committee Again

New report from the UK Sci Tech Committee. (I’m traveling – see Bishop Hill for link.) My take is that the Committee was annoyed with the University of East Anglia, being quite critical of the inquiries in the running text, but have decided that there are other more pressing priorities and that it’s time to “move on”.
In some cases, they seem to have gritted their teeth and accepted untrue statements at face value. Graham Stringer, by far the most knowledgeable member of the Committee on matters UEA, moved a critical amendment to the conclusions that is an honest appraisal of the situation.

Continue reading →

Jeff Id

I’m sorry to learn that Jeff Id has suspended operation of his blog in order to properly carry out his obligations to his business and his young family.

Jeff introduced himself to Climate Audit soon after he started his blog (here)>. He began with a variety of interesting technical analyses of Mann et al 2008 – technical analyses of the type that interested me and CA readers here. My first mention of the blog was here.

I was a regular reader and will miss my daily visit to his blog. Jeff plans to stay in touch.

I don’t know how he managed to balance the responsibilities of a business and a young family with blog activity as long as he did, but am grateful that he managed as long as he did.

Was Phil Jones an IPCC Virgin?

A few days ago, I challenged Trenberth’s claim that “AR4 was the first time Jones was on the writing team of an IPCC Assessment.”

Earlier this year, Real Climate stated that AR4 had been “written by over 450 lead authors and 800 contributing authors”. In my challenge to Trenberth’s claim, I observed that Jones had been a Contributing Author to the 2001 and 1995 IPCC Assessment Reports (Pielke Jr later adding that Jones had been a Contributing Author to the 1990 Assessment Report.) Ergo, Trenberth’s claim that AR4 was the “first time Jones was on the writing team of an IPCC Assessment” was untrue.

To most people, that would end the discussion about whether AR4 had been Jones’ first time or not.

However, Dave Clarke aka Deep Climate has now argued that “contributing authors are not on the [IPCC] writing team” and that

Trenberth makes it crystal clear that he is means that Jones was a “first time” lead author.

Clarke’s idea that contributing authors are not part of the IPCC writing team will no doubt come as a surprise to realclimate – who are, no doubt, scrambling as we speak to correct their previous mis-statements on this point.

In addition, Jones was not merely a “Contributing Author” to AR3. Jones was part of the writing team for AR3 Chapter 3 – described in IPCC email as a “Key Contributor”. The term “Key Contributor” is not used in IPCC documents, but was used to describe the role of Jones and several others in the preparation of AR3 Chapter 2, where Jones was assigned responsibility for writing part of AR3 Chapter 2. The term was used in an IPCC email of June 21, 1999 (929985154.txt in the Climategate dossier) with Jones an addressee (but not Trenberth). (In the eventual listing of Chapter 2 authors, the Key Contributors are listed ahead of “ordinary” Contributors Authors.)

The online version of this Climategate email is truncated for some reason. It shows only the following:

Below is the text and attached is a file in MSWord regarding a plan of
action for Chapter 2 leading up to the IPCC Meeting in Arusha, Tanzania.

June 21, 1999

Dear Lead Authors and Key Contributors,

This note is to outline a plan of action for Chapter 2 leading up to the
IPCC meeting in Arusha, Tanzania to take place 1-3 September. As you know,
we are now in the midst of a

The complete email clearly shows Jones’ involvement in the writing process:

From: sdecotii@
To: christy@, clarkea@, @cabel.net, pfrich@, pgroisma@, jwhurrell@,
m.hulme@, p.jones@, Jouzel@, mann@, j.oerlemans@, deparker@,tpeterso@, drind@, drobins@,j.salinger@, walsh@, swwang@

Subject: Plan of action for Chapter 2
Date: Mon, 21 Jun 1999 13:12:34 -0400
Below is the text and attached is a file in MSWord regarding a plan of
action for Chapter 2 leading up to the IPCC Meeting in Arusha, Tanzania.

June 21, 1999
Dear Lead Authors and Key Contributors,
This note is to outline a plan of action for Chapter 2 leading up to the
IPCC meeting in Arusha, Tanzania to take place 1-3 September. As you know,
we are now in the midst of a
“friendly review” from our colleagues of the
strawman draft of our chapter. We expect to receive comments from these
reviews through middle or even late July. These reviews will include some
from people other than our nominated reviewers, like Sir John Houghton,
from whom we have just had a brief review. Please check regularly with the
Tar02.meto.gov.uk email site to cover this aspect.

Accordingly we ask each of the individuals listed below to revise the draft
section as suggested below, and to indicate their response to reviewer’s
comments. The first person listed is to take the lead, and individuals
with an asterisk by his name are to prepare the material for presentation
in Arusha. We would ask that a provisionally revised part of your chapter
be completed by 20 August and emailed to Tom Karl or placed on the web-site
so that Sylvia Decotiis can create a new version of Chapter 2 for Tom to
bring to Tanzania. Tom will bring one paper copy of the provisional new
“Arusha” version of chapter 2 to Tanzania, and a complete series of
electronic files which can be input to PCs via 1.4MB floppy disks. It would
be a considerable advantage for attendees to bring portable PCs, though we
expect some IPCC PCs to be available at the Arusha International Conference
Centre.

Chris Folland will be leaving for Tanzania early (24 Aug) whereas Tom Karl
will still be available until 29 Aug for urgent interactions. We will
decide later as to whom, and how many of us, should actually make
presentations, noting that Hans Oerlemans is not likely to be present. But
all attendees be prepared, and bring appropriate visual material and of
course, further suggestions. We have listed assignments next to each
section.

Section 2 —– Tom Karl* and Chris Folland* Executive Summary — total
revision and update
Section 2.1 —- Chris Folland* Changes needed regarding uncertainty
guidelines
Section 2.2.1 —- Chris Folland* Okay for now
Section 2.2.2 —- David Parker, Phil Jones, Tom Peterson, Chris Folland*
Length okay, but reduce number of figures.

Section 2.2.3 —- John Christy* Check for accuracy
Section 2.2.4 —- John Christy* Check for accuracy
Section 2.2.5 to 2.2.6 —- Oelermans*, Nick Rayner, John Walsh, David
Robinson, Tom Karl and Chris Folland. Glacier section needs to be updated
Section 2.2.7 —- Oelermans, Tom Karl* Check for accuracy
Sections 2.3 through Section 2.3.5—- Mike Mann*, Phil Jones Reduce in
size by about 10%

Section 2.4 through Section 2.4.5 —-Jean Jouzel* Reduce in size about 10%
Section 2.5 through 2.5.4 —- Jim Salinger*, Pasha Groisman, Mike Hulme,
Wang. Provide a better context for why this section is important, more on
upper tropospheric water vapor if possible
Section 2.5.5 —- Steve Warren, Dale Kaiser, Tom Karl* Add new analyses of
cloud amount
Section 2.5.6 —-Jim Salinger*
Section 2.6 through 2.6.6 —-Jim Salinger*, George Gruza, Alynn Clarke,
Wang. Reduce in size by at least 50%. Identify a rationale section at the
beginning. IPCC 1995 will help here. Some material may go elsewhere. May
need to consult Mike Mann or Jean Jouzel. Please send revised section to
Chris Folland to finally review (even if not complete) by 16 August. Chris
will feed back changes to Jim by 23 August. Jim Salinger should interact
with Chris during this work too. Jim should prepare presentational material
Section 2.7 through 2.7.4 —-David Easterling, Pasha Groisman, Tom Karl*

Review for accuracy
Povl Frich: please interact and be prepared to present extremes parts. Jim
Salinger: you may have more material on extremes in the South Pacific.
Please feed this to Tom Karl and Povl Frich.
Section 2.8 —- Tom Karl, Chris Folland* Develop a summary, including
strawman cartoon
In addition we have about twice the number of figures that will be allowed
so everyone should identify figures that can be removed or combined to
reduce the size. The latter can sometimes be very effective. At the
present time we are about 1/3 over our word limit so everyone will have to
respond to the reviewers (often requesting more), and yet being more
judicious in the words we use. Please consult the 1995 IPCC Report as a
guide.

Please do not hesitate to comment on these plans, preferably as soon as
possible, so that holiday arrangements etc do not cause problems.
Cheers and thanks,
Chris and Tom

(See attached file: ARUSHA INSTR LEAD AUTHORS.doc)
Attachment Converted: “c:\eudora\attach\ARUSHA INSTR LEAD AUTHORS.doc”

The document “Arusha Instr[uctions?] Lead Authors.doc” is not in the Climategate documents. However, Jones received this document, which presumably set out the duties of Lead Authors (and Key Contributors).

And, of course, following the Arusha meeting, Jones was intimately involved in correspondence with Mann, Briffa and Folland about what to do about the Briffa reconstruction – correspondence that led on the one hand to the deletion of post-1960 data in the IPCC graphic and on the other hand to the notorious ‘hide the decline’ email about the WMO graphic.

Clarke also consulted Trenberth’s CV and observes that Trenberth’s offices in previous IPCC reports had been senior than Jones’. Be that as it may, that doesn’t make Jones an IPCC virgin.

In IPCC’s public face, Contributing Authors are regularly counted as part of the IPCC writing team. Plus, in Jones’ individual case, although he was “only” an AR3 contributing author, he was nonetheless considered a “Key Contributor” and had been actively involved as part of the Chapter 2 writing team. Trenberth’s statement that AR4 was the “first time Jones was on the writing team of an IPCC Assessment” was untrue on either count.

Team Policy on Acknowledgements

After CA reported Trenberth’s lifting of text from Hasselmann 2010 verbatim or near-verbatim either without citation or, in the one citation, a citation that was inadequate given the lengthy near-verbatim quotation, Trenberth moved quickly to cooper up his presentation against plagiarism allegations by inserting citations to Hasselmann 2010, responding to each of the incidents reported at CA. Trenberth did not acknowledge Climate Audit.

Question: given that Trenberth considered the problems sufficient to justify making changes, should Trenberth have acknowledged Climate Audit for drawing the problem to his attention? Continue reading →

Trenberth and Lifting Text Verbatim #2

On January 14, 2011, I reported here that Trenberth’s AMS presentation had lifted text verbatim or near-verbatim from Hasselmann 2010 with no citation in most cases and, in the one case where Hasselmann 2010 was cited, the citation was insufficient under standard academic practices given the lengthy near-quotation. Trenberth’s original presentation is here.

This post has obviously been brought to Trenberth and/or AMS’s attention, as they have deleted the original version of Trenberth’s presentation and replaced it with an amended version, without a change notice.

The amended version picks up most of the problems raised in the previous CA post. Here are the points raised in the CA post and Trenberth’s changes:

Trenberth originally stated:

Scientists make mistakes and often make assumptions that limit the validity of their results. They regularly argue with colleagues who arrive at different conclusions. These debates follow the normal procedure of scientific inquiry.

The amended version:

Hasselmann (2010) further notes that scientists make mistakes and often make assumptions that limit the validity of their results. They regularly argue with colleagues who arrive at different conclusions. These debates follow the normal procedure of scientific inquiry.

Trenberth’s originally statement about tactics to use against “deniers”:

It is important that climate scientists learn how to counter the distracting strategies of deniers. Debating them about the science is not an approach that is recommended.

The amended version:

It is important that climate scientists learn how to counter the distracting strategies of deniers (Hasselmann 2010). Debating them about the science is not an approach that is recommended.

Trenberth originally stated:

The main societal motivation of climate scientists is to understand the dynamics of the climate system (both natural and human induced), and to communicate this understanding to the public and governments.

The amended version:

The main societal motivation of climate scientists is to understand the dynamics of the climate system (both natural and human induced), and to communicate this understanding to the public and governments (Hasselmann 2010).

Trenberth did not feel obligated to restate everything that Hasselmann had stated. For example, Trenberth did not repeat Hasselmann’s observation that:

Individually, most climate scientists have the goal of establishing a scientific reputation and, if possible, attaining more public funding for climate research.

Trenberth originally stated:

They [climate scientists] have faith in the scientific method and the efficacy of the established peer-review process in separating verifiable scientific results from baseless assertions.

The amended version:

As Hasselmann (2010) further notes, they have faith in the scientific method and the established peer-review process in separating verifiable scientific results from baseless assertions.

As to the lengthy introductory paragraph which was lifted near-verbatim from Hasselmann, but with no indication that large sections were verbatim: the original Trenberth version was:

Three investigations of the alleged scientific misconduct of the Climate Research Unit at the University of East Anglia — one by the UK House of Commons Science and Technology Committee, a second by the Scientific Assessment Panel of the Royal Society, chaired by Lord Oxburgh, and the latest by the Independent Climate Change E-mails Review, chaired by Sir Muir Russell — have confirmed what climate scientists have never seriously doubted: established scientists depend on their credibility and have no motivation in purposely misleading the public and their colleagues. Moreover, they are unlikely to make false claims that other colleagues can readily show to be incorrect. They are also understandably (but inadvisably) reluctant to share complex data sets with non-experts that they perceive as charlatans (Hasselman 2010)

The amended version is:

As noted by Hasselmann (2010), three investigations of the alleged scientific misconduct of the Climate Research Unit at the University of East Anglia — by the UK House of Commons Science and Technology Committee and by the Lord Oxburgh and Sir Muir Russell reviews — have confirmed that established scientists depend on their credibility and have no motivation in purposely misleading the public and their colleagues. Moreover, they are unlikely to make false claims that other colleagues can readily show to be incorrect. They are also understandably (but inadvisably) reluctant to share complex data sets with non-experts that they perceive as charlatans (Hasselman 2010).

Note here that Trenberth removed the absurd Hasselmann characterization of the Oxburgh inquiry that had passed muster with the editors of Nature Geoscience – that it was the “Scientific Assessment Panel of the Royal Society”. A point also raised in comments in the CA post.

Trenberth did not submit a comment to Climate Audit thanking us for enabling him to mitigate the problem prior to the actual formal presentation of his speech or otherwise thank us at the AMS webpage at which the changes were made.

Even as amended, Trenberth’s use of extended near-quotations without quotation marks must surely still be at the edge of acceptable practice, if not over it.

In appraising whether inadequate citation rises to being academic misconduct, it seems to me that one needs to consider whether there is a claim or implied claim of originality and how such incidents are handled in the field. Trenberth’s speech was not the same thing as a master’s thesis. Trenberth had lifted text, but AMS appears to have decided that the situation could be more or less coopered up by providing more citations to Hasselmann; that the inadequacy of the citations did not entail that the entire speech be withdrawn; or that AMS was obliged to file a complaint against Trenberth for academic misconduct.

It’s also interesting what Trenberth chose to change and not to change. Plagiarism is an issue that is uniquely central to the academic world and Trenberth moved quickly to erase any evidence of plagiarism. The non-academic world would be less concerned about plagiarism and more concerned about Trenberth’s use of the offensive term “denier” and whether Trenberth’s Empire Strikes Back attitudes are a useful contribution to the post-Climategate debate. Although Trenberth was also criticized on these counts, Trenberth made no concessions or changes to this aspect of his speech.

Trenberth and Lifting Text Verbatim

In case readers think that Trenberth’s outburst discussed yesterday represents an isolated and unfortunate climate scientist incident, this is not the case. In fact, some of Trenberth’s most objectionable language was lifted verbatim from an article in Nature Geoscience earlier this year. Trenberth here; Hasselmann here.

Trenberth’s copying from Hasselmann came in two forms:
– Trenberth copied one long paragraph verbatim mostly verbatim without quotation marks. While Hasselmann was cited at the end of the paragraph, the fact that the text was lifted [mostly] verbatim was not shown – something that John Mashey will no doubt weigh in on.
– second, Trenberth copied multiple sections of Hasselmann either verbatim or with negligible paraphrase without any citation whatever. Continue reading →

Trenberth’s Bile

Anthony draws attention to a bilious diatribe by Trenberth against “deniers”.

I have some back-history with Trenberth. In 2005, Trenberth was interviewed by Paul Thacker of ES&T about the MM articles (discussed here) where he stated:

There have been several examples of people who have come into the field of climate change and done incredibly stupid things by applying statistics in ways that are inappropriate for the data, [Trenberth] says.

I wrote back and forth with Trenberth a number of times in respect to his earlier comments about me – the correspondence is online here. After several attempts to get Trenberth to justify his allegations, Trenberth challenged me to respond to the criticism at realclimate. When I did so, Trenberth discontinued the correspondence without justifying his comment.

In his most recent outburst, Trenberth says:

Debating them [“deniers”] about the science is not an approach that is recommended. In a debate it is impossible to counter lies, and caveated statements show up poorly against loudly proclaimed confident statements that often have little or no basis. Scientific facts are not open to debate and opinion because they are evidence and/or physically based. Moreover a debate actually gives alternative views credibility.

Trenberth described his recommended tactic in a Climategate email as follows:

So my feeble suggestion is to indeed cast aspersions on their motives and throw in some counter rhetoric. Labeling them as lazy with nothing better to do seems like a good thing to do.

Trenberth now complains that the supposedly “false claims” of critics have not been “scrutinized or criticized” enough:

But their critics are another matter entirely, and their false claims have not been scrutinized or criticized anything like enough!

However, Trenberth himself advocated the strategy of casting aspersions on critics instead of scrutinizing their arguments and, as one of the architects of this strategy, is hardly in a position to complain.

The way to counter lies is obvious – show evidence that statements are lies. For example, when Mann said that I had asked for an Excel spreadsheet and that they had inadvertently introduced errors in the process of tailoring the data for this special request, the way to counter it is to produce the original email showing that we had not asked for an Excel spreadsheet but an FTP location and that the dataset that we were directed to at Mann’s FTP site was dated long prior to my inquiry. (The data set was deleted by Mann shortly after this incident, thereby removing this evidence.) Or when Mann told the NAS panel that he hadn’t calculated a verification r2 statistic as this would be a “foolish and incorrect” thing to do, the way to counter this was to examine his original article which showed the verification r2 statistic for the AD1820 step and, when code became available for this step, to show that the code calculated the verification r2 statistic in the same step as the RE statistic that was reported.

Trenberth also purports to justify Jones’ successful effort to keep McKitrick Michaels 2004 out of the two AR4 drafts sent to reviewers on the basis (this incident has been discussed at length on other occasions) that:

AR4 was the first time Jones was on the writing team of an IPCC Assessment.

while noting that Trenberth himself, as a “veteran”, was aware of the obligations:

As a veteran of 3 previous IPCC assessments I was well aware that we do not keep any papers out, and none were kept out.

Trenberth goes on to add that:

[climate scientists] are unlikely to make false claims that other colleagues can readily show to be incorrect.

Unfortunately, we’ve seen too many incidents where climate scientists make false claims that are readily shown to be incorrect. We need think back no further than Jones’ claim that CRU had confidentiality agreements that contained language prohibiting the distribution of data sent to Peter Webster to “non-academics”.

Trenberth’s very claim that AR4 was the first time that Jones had been on a writing team is itself another example of an untrue statement that can be “readily” demonstrated to be untrue (although his “colleagues” have thus far not called him on it.)

Both Jones and Trenberth are listed as contributing authors of AR3. (Indeed, Jones’ correspondence about the Briffa reconstruction in the wake of the 1999 Lead Authors meeting in Arusha, Tanzania was important in the setting of the notorious “hide the decline” memo.) See the list of AR3 chapter 2 authors below, where both Trenberth and Jones are listed as Contributing Authors.

Likewise with AR2 – both Trenberth and Jones had precisely the same standing as AR2 Contributing Authors. See here.

Trenberth’s assertion that Phil Jones had not previous been on an IPCC writing team is, to borrow a phrase, a “travesty”. It turns out that Phil Jones had been involved in all three previous IPCC reports.

Update: Pielke Jr observes:

FYI, Phil Jones was also listed as a “Contributor” to IPCC FAR (1990) Sections 6, 7 and 8 pp. 348-349. (Later this role was called ‘Contributing Author”).

[Update Jan 20] Trenberth’s CV (as pointed out by Dave Clarke) shows that Trenberth had also been a Lead Author in 1995 and 2001. However, as discussed here, Jones had not merely been a Contributing Author to IPCC 2001, he had been a Key Contributor who had been part of the Chapter 2 “writing team”.

The Hockey Stick in Art and Literature

One of Josh’s finest graphic flourishes: (h/t Bishop Hill and WUWT)

UC on Mannian Smoothing

Two comments from UC on smoothing CET using Mannian smoothing, a technique peer reviewed by real climate scientists (though not statisticians).

I think these coldish years do matter, maybe now there will be some advance in smoothing methods. mike writes (1062784268.txt) ( I think this is somewhat related to CET smoothing (?) )

The second, which he calls “reflecting the data across the endpoints”, is the constraint I have been employing which, again, is mathematically equivalent to insuring a point of inflection at the boundary. This is the preferable constraint for non-stationary mean processes, and we are, I assert, on very solid ground (preferable ground in fact) in employing this boundary constraint for series with trends…
mike

I assert that a preferable alternative, when there is a trend in the series extending through the boundary is to reflect both about the time axis and the amplitude axis (where the reflection is with respect to the y value of the final data point). This insures a point of inflection to the smooth at the boundary, and is essentially what the method I’m employing does (I simply reflect the trend but not the variability about the trend–they are almost the same)…

And now this leads to following figure:

mr2010

Jones also mentions CET:

Normal people in the UK think the weather is cold and the summer is lousy, but the CET is on course for another very warm year. Warmth in spring doesn’t seem to count in most people’s minds when it comes to warming.

And later here:

Yes, extrapolations are problematic if someone bothers to check those later:

c1

Can’t understand what Mann means by ‘preferable constraint for non-stationary mean
processes’.. I’d prefer no smoothing at all if there is no statistical model for the process itself, something like this maybe:

c2

Update (UC, 8 Jan 2011)

Code in here .

For CA readers it is clear why Minimum Roughness acts this way, see for example RomanM’s comment in here (some figures are missing there, will try to update). But to me it seems that the methods used in climate science evolve whenever temperatures turn down (Rahmstorf example is here somewhere, and you can ask JeanS what has happened in Finnish mean temperature smooths lately).