Abstract. What sets the boundary of the safety net? In a crisis a central bank rescues some private near-monies and lets others fail, and a tempting answer to which is that rescue tracks an entity’s position in the payment, settlement and collateral plumbing (its “circulation-centrality”) rather than its size or its legal category. This paper builds that hypothesis into a falsifiable test and reports that it fails. On a panel of thirty rescue decisions from 1970 to 2023, a blind-coded index of circulation-centrality at first appears to separate the rescued from the abandoned. The appearance does not survive scrutiny. Blind coding removes a thirteen-point hindsight inflation in the author’s own scoring; dropping two index components that quietly restate the rescue trigger removes the construct’s circular content; and correcting two contested keystone cases, Lehman Brothers (whom the Federal Reserve could have saved, per Ball 2018) and Washington Mutual (whose depositors were in fact protected), removes the rest. Conditioning on legal-political capacity rather than deleting the inconvenient cases, the circulation signal falls to a partial correlation of 0.13 (p = 0.53) and is out-predicted by a plain systemic-importance indicator. When an independent coder, blind to the hypothesis, reconstructs the capacity and mandate covariates, the signal does not appear at all (partial r = −0.26, not significant). The boundary of last-resort lending is governed by size and political-legal capacity, the familiar too-big-to-fail account, not by a distinct circulation construct. We offer the test as a reusable template and the false positive as a cautionary anatomy: a strong-looking result manufactured by circular proxies, conditioning-by-deletion, and two miscoded historical cases.

Keywords: lender of last resort, too big to fail, systemic importance, central-bank rescue, financial safety net, blind coding, replication, null result.

JEL: E58, G01, G28, N20, B41.


1. Introduction

In September 2008 the Federal Reserve lent against collateral it had never accepted, to firms it did not supervise, and opened swap lines so that banks abroad could fund dollar books no American sovereign had authored. It saved some institutions and let others fail. The question this paper asks is the one those choices pose: what decides which? Lehman Brothers, at the centre of the repo system, was allowed to fail; a $5bn bank in 1974 was saved. A theory of the modern safety net has to say what the rescues have in common and the abandonments lack.

One answer has been gathering momentum in the money-view and shadow-banking literatures. On this view the safety net follows function, not form: once near-banking migrates outside the chartered perimeter, the central bank can no longer confine its protection to a legal category, so it rescues whatever has become central to the system’s plumbing (Schwarcz 2010; Nersisyan 2015; Buiter et al. 2023). Push that intuition one step further and it becomes a sharp, testable claim, which I will call the circulation criterion: whether an entity is rescued tracks its circulation-centrality (its position in the payment, settlement and collateral chains through which the host’s money moves) rather than its size or its legal category. A near-money, throughout, is a private claim that circulates as if it were central-bank money, redeemable at par and accepted with no questions asked, while being a claim on someone else.

The circulation criterion is attractive. It promises to explain the post-2008 expansion of last-resort lending without appeal to politics or to crude size, and it gives the money-view’s “function not form” slogan an operational edge. This paper takes the claim seriously enough to test it, and finds that it does not hold. On a panel of thirty rescue decisions spanning 1970 to 2023, a carefully built, blind-coded index of circulation-centrality at first seems to separate the rescued from the abandoned. That separation turns out to be an artifact of three things, each of which I can isolate and remove: hindsight in the author’s own coding, two index components that secretly restate the very outcome they are meant to predict, and two historical cases miscoded in ways that favour the hypothesis. Strip all three, keep every case in the analysis rather than discarding the awkward ones, and the circulation signal vanishes. What remains is the familiar story: rescue is predicted by plain systemic importance and by the legal and political capacity to act, the too-big-to-fail and political-economy account that Bagehot, Kane and Ball would each recognise.

The contribution is therefore a disciplined negative, and a method for reaching it. I claim three things, and only the first is about central banking. First, the circulation criterion, as operationalised here and tested on this panel, does not out-predict systemic importance; the boundary of the safety net is not a distinct “plumbing” phenomenon. Second, the apparent support for it had a recoverable anatomy (construct circularity, conditioning-by-deletion, and contested case-coding) that recurs across the empirical study of rare, high-stakes decisions, and the paper is a worked specimen of how each manufactures a false positive. Third, the protocol that exposed the artifact (blind coding of the predictor, covariates entered as controls rather than as filters, and adversarial fact-checking of the keystone cases) is a reusable template for testing rescue-boundary hypotheses, where samples are small, outcomes are famous, and the temptation to fit a clean story is acute.

A reader may object that a null result on thirty hand-coded decisions settles little. That is fair, and Section 7 states the limits plainly: this is a discriminating-case design, the statistics are illustrative rather than powered, and the finding is about one operationalisation of one construct. But the value of the exercise does not rest on statistical power. It rests on showing that a result which looked powered and clean, a point-biserial correlation of 0.63 with the trappings of pre-registration and blind coding, dissolves entirely under three disciplines that any referee can demand, and that the dissolving is itself informative about why we should distrust such results.

The paper proceeds as follows. Section 2 states the circulation hypothesis at its strongest. Section 3 describes the panel and the three disciplines the test applies. Section 4 reports the staged collapse of the apparent result. Section 5 identifies what does predict rescue. Section 6 sets out the anatomy of the false positive as a methodological contribution. Section 7 gives the implications and the limits. Section 8 concludes.

2. The hypothesis under test

It is only fair to a hypothesis to state it at its strongest before testing it, so that a failure is a failure of the idea and not of a strawman. The circulation criterion has an evocative motivation, which I will give and then set aside.

The motivation is a biological analogy. Peripheral near-monies can be read as obligate symbionts on the central-bank “host”: they circulate as dollars by borrowing the host’s organs, its unit of account, its capacity to settle with finality and its backstop, because, by the social ontology of money, they cannot author their own (Searle 1995; Aglietta and Orléan 1998; Mehrling 2011). A rescue, on this reading, is the host drawing a symbiont inside its membrane when the symbiont’s failure would seize up the host’s own circulation. The analogy yields a boundary: the host absorbs what would damage its circulation, and ignores what would not. Whatever one makes of the metaphor, it points at a concrete and testable prediction, and the prediction is all that matters here.

The circulation criterion (the hypothesis under test). In a systemic crisis, whether an entity is backstopped tracks its centrality in monetary circulation, not its institutional category or its size, once all three are measured.

Two features make this a genuine hypothesis rather than a restatement of consensus. It is directional: more circulation-central entities should be more likely to be rescued, holding category and size fixed. And it is discriminating: it predicts that where circulation-centrality and size point in different directions, circulation should win. A $5bn bank that sits at a settlement chokepoint should be rescued; a $300bn retail bank that does not should be let go. If the prediction held, it would sharpen the money-view’s “function not form” intuition into something with teeth, and it would distinguish the boundary of the safety net from the ordinary “too big to fail” account, on which rescue tracks size and interconnectedness as such.

The rival the criterion must beat is therefore not a vacuum but a well-developed account: that rescue tracks systemic importance in the conventional sense (size, leverage, interconnectedness) and the legal and political capacity to act (Kane 1981; Acharya and Yorulmazer 2007; and, on the decisive Lehman case, Ball 2018). The whole question is whether “circulation-centrality” names something this rival does not already capture.

3. Data and method

3.1 The panel

The unit of observation is a rescue decision: an episode in which the authorities faced a distressed near-money or near-money class and chose whether, and how far, to draw it inside the safety net. The panel holds thirty such decisions across eight episodes: the 1970 Penn Central commercial-paper crisis, the 1974 Franklin National failure, the 1984 Continental Illinois rescue, the 1990 Drexel Burnham failure, the 1998 Long-Term Capital Management episode, the crises of 2008 and 2020, and the 2023 regional-bank and stablecoin episode. Decisions are coded at the level at which the intervention was pitched, which is mixed: sometimes a named firm, sometimes a class or market. The sample was built to include the cases that could refute the criterion as well as confirm it, among them entities that were systemically adjacent and not rescued (Penn Central, Drexel, Lehman) and one rescued by private coordination rather than public money (LTCM).

This is a discriminating-case design, not a powered panel. Thirty decisions will not support delicate inference, and the statistics below are reported as illustrative throughout. The argument does not turn on a p-value; it turns on what happens to an apparent result as successive disciplines are applied.

3.2 The predictor and the outcome

The outcome is an ordinal rescue-intensity from 0 (none) to 4 (direct capital or bespoke rescue), collapsed to a binary rescued indicator at intensity of one or more.

The predictor is a Circulation-Centrality Index built from six proxies, each scored 0 (absent) to 3 (core): position in large-value payment systems; role in securities and FX settlement; centrality in repo and collateral intermediation; the money-likeness of the entity’s own liabilities; gross funding interconnectedness with core banks and dealers; and the non-substitutability of its function at short notice. The six are summed and rescaled to 0–100. The competing predictors are legal category, raw size (assets), and a conventional systemic-importance indicator (G-SIB or SIFI status, with an implicit-large flag before 2011).

3.3 Three disciplines

The test applies three disciplines, and the result is read through what each does to the apparent correlation.

Blind coding of the predictor. Each entity was reduced to a neutral pre-crisis profile with its name and its rescue outcome removed. Two coders, given only the rubric and told no outcomes, scored the six proxies independently. Inter-coder agreement and the gap between the blind index and the author’s own sighted coding are both reported, the second because it measures hindsight directly.

Covariates as controls, not filters. The host is not a unitary optimiser. Officials act under a legal and collateral capacity to act and, sometimes, under a pre-existing fiscal mandate (a Congressional or Treasury authority such as the CARES Act). An earlier version of this analysis used those two covariates to select a “clean” subset, discarding every case in which capacity was absent or a mandate intruded. That is conditioning by deletion, and it is the wrong move, because the discarded cases are exactly the ones that discipline the claim. Here the covariates are entered as controls in a model on all thirty cases, and the reported statistic is the partial correlation of rescue with circulation-centrality, holding capacity, mandate and systemic importance fixed. To close the last channel by which the author’s hand could shape the result, the capacity and mandate codes were themselves reconstructed by an independent coder, told the rubric but not the hypothesis (Section 4).

Adversarial fact-checking of the keystones. Two cases carry disproportionate weight: Lehman, the highest-centrality entity that was not rescued, and Washington Mutual, the largest bank that “failed”. Each is checked against the documentary record rather than coded from the author’s prior, and where the record is contested the case is coded against the hypothesis to see whether the result survives.

4. Results: the staged collapse

The apparent result is real, and it is worth seeing before it falls. On a favourable reading (the author’s own sighted coding, the full six-proxy index, and the clean subset carved out by deletion), circulation-centrality separates the rescued from the abandoned at a point-biserial correlation of 0.63 (p = 0.016), with rescue rising monotonically across three tiers. Taken at face value, that is a finding. It is not one, and the three disciplines show why.

Discipline 1: blind coding catches a large hindsight effect. The two blind coders agree almost perfectly with each other (89% of 180 cells exactly, the rest within a single point), so the rubric is reliable. But the blind index sits about thirteen points below the author’s own sighted scores, and the gap is systematic: the entities that turned out to be central had been scored high by a coder who knew the history. The construct is reliable between coders and yet contaminated by hindsight when the coder can see outcomes. Reliability is not validity. All figures below use the blind index.

Discipline 2: two proxies restate the outcome. Four of the six proxies are mechanical facts about plumbing. Two are not. “Non-substitutability at short notice” and “money-likeness of liabilities” are close paraphrases of “this entity’s failure would seize the system”, which is the rescue trigger itself. An index that contains a near-synonym of the outcome will predict the outcome by construction. Rebuilding the index from the four genuinely mechanical proxies (a plumbing-only index) is the conservative move. On the full sample the plumbing-only index correlates with rescue at just 0.08, and it is out-predicted by both the systemic-importance indicator (0.22) and size (0.20). The headline “circulation beats size and category” fails at the simplest comparison once the index cannot lean on the two outcome-laden components.

Discipline 3: keeping the cases, and correcting two of them, removes what is left. Entered as controls rather than filters, the covariates still leave an apparent partial correlation of 0.42 (p = 0.029) under the author’s coding. That residual does not survive the documentary record.

  • Lehman. The paper had coded Lehman as a non-rescue forced by the absence of capacity, citing the Fed’s own account that it lacked sufficient collateral. Ball’s (2018) book-length study of exactly this question concludes the opposite: the Federal Reserve had the legal authority and Lehman had adequate collateral, and the decision to let it fail was political, taken primarily by the Treasury Secretary. Coding Lehman’s capacity as present, per the standard scholarly treatment, drops the partial correlation to 0.34 (p = 0.078), no longer significant.
  • Washington Mutual. The paper had coded WaMu as the flagship “abandonment”, a $300bn bank let go. But the FDIC resolved WaMu through a purchase-and-assumption sale to JPMorgan in which all depositors, insured and uninsured, were protected at no loss to the insurance fund; only the holding company’s investors bore losses. By the criterion’s own logic, in which a protected near-money is a rescued one, WaMu’s deposits were rescued. Recoding it accordingly drops the partial correlation to 0.21 (p = 0.29).

Applying both corrections together, which the record requires, leaves a partial correlation of 0.13 (p = 0.53), indistinguishable from zero. The staged collapse is summarised in Table 1.

Table 1. The circulation signal under successive disciplines (N = 30).

Specificationpartial r (rescued, circulation | systemic-importance, capacity, fiscal)
Author coding, plumbing-only index0.42 (p = 0.029)
+ Lehman capacity present (Ball 2018)0.34 (p = 0.078)
+ Washington Mutual depositors protected0.21 (p = 0.29)
+ both corrections (required by the record)0.13 (p = 0.53)
+ independent, hypothesis-blind coding of the covariates−0.26 (p = 0.19)

A final discipline removes any lingering worry that the covariates were coded to fit. An independent coder, given the rubric but not the hypothesis, reconstructed capacity and fiscal mandate for all thirty cases, disagreeing with the author’s coding on seven and eight cases respectively. Under that independent coding the partial correlation is −0.26 (p = 0.19): not merely absent but, if anything, pointing the other way, because the reshuffled set now places Lehman inside the test as a high-centrality non-rescue and Washington Mutual as a low-centrality rescue. Whatever the apparent result was, it does not survive a coder who cannot see the hypothesis.

The strong first impression, then, was the joint product of three things: a coder who could see outcomes, two proxies that restated the outcome, and two keystone cases coded in the hypothesis’s favour against the documentary record. Remove all three and the circulation criterion has no purchase the rival account does not already have.

5. What does predict rescue

If circulation-centrality is not doing the work, what is? Two things the rival account names.

The first is conventional systemic importance. The systemic-importance indicator out-predicts the plumbing-only circulation index univariately (0.22 against 0.08), and size is close behind (0.20). The entities the central bank caught were, by and large, the large and interconnected ones. This is the too-big-to-fail account, and the data do not improve on it.

The second is legal and political capacity. The cases the circulation story had to explain away are, on the record, decided by capacity and politics. Lehman was not a low-centrality entity that fell below a plumbing threshold; it was a high-centrality entity that the Fed could have saved and, for political reasons, did not (Ball 2018). The broad 2020 and 2023 interventions were not circulation-graded; they followed explicit fiscal mandates. Once these are read as what they are, the residual variation that a “circulation” construct might have explained is small and statistically silent.

The honest synthesis is unglamorous and was available before the exercise began. The boundary of the safety net is set by how big and interconnected an entity is, and by whether the authorities have the legal room and the political will to act. That is Bagehot brought up to date by Kane and corrected, on the pivotal case, by Ball. The contribution of this paper is not to discover that. It is to show, with a disciplined test, that the more interesting alternative does not survive contact with the evidence.

6. The anatomy of a false positive

The methodological value of the exercise is that the apparent result failed in three identifiable ways, each of which is a general hazard in the empirical study of rare, high-stakes decisions. Naming them is the paper’s transferable lesson.

Construct circularity. When a predictor is hand-built from components, it is easy to include, often without noticing, a component that paraphrases the outcome. “Would its failure be catastrophic?” is not a clean input to a model of catastrophe-avoidance; it is the thing to be explained. The tell is that the suspect components are the ones doing the predictive work: here, removing two of six proxies took the univariate correlation from the high-CCI story to 0.08. The discipline is to separate, in advance, the mechanically observable inputs from any component that encodes a judgement about consequences, and to report the result on the former alone.

Conditioning by deletion. Faced with cases that contradict a hypothesis, an analyst can introduce a covariate that “explains” them and then restrict the test to the cases the covariate leaves behind. This feels like control and is in fact selection. The disconfirming cases carry most of the information, and discarding them all but guarantees a clean residual. The discipline is to enter such covariates as controls in a model that keeps every case, so that the awkward observations push back on the estimate rather than disappearing from it. Here that single change took a significant clean-subset result to an insignificant full-sample one.

Contested case-coding by the interested party. In a small panel of famous events, one or two cases are load-bearing, and they are often exactly the cases on which the historiography is divided. An author who codes them from a prior, rather than from the contested record, can move the result without intending to. The Lehman capacity code is the example: a single contested assignment, taken from the rescuer’s own account and against the standard scholarly treatment, was the difference between a refuting case and an excluded one. The discipline is to identify the load-bearing cases, code them against the hypothesis where the record is contested, and report the sensitivity.

None of these is exotic, and that is the point. Each is a routine way for a strong-looking result to be manufactured, and the three together produced, in this instance, a point-biserial of 0.63 that was indistinguishable from zero once they were undone.

7. Implications and limits

The substantive implication is modest and was anticipated by the literature the paper set out to sharpen. The boundary of last-resort lending is not a distinct “circulation” phenomenon; it is governed by systemic importance and political-legal capacity. For policy this points back to the unfinished too-big-to-fail agenda rather than towards a new “plumbing-centrality” target: the entities a central bank will catch are the large and interconnected ones it has the room and the will to catch, which is what the resolution and capital literatures already address.

The methodological implication is a reusable protocol for testing rescue-boundary hypotheses, where the obstacles are generic: small samples, famous outcomes, and hand-built predictors. Blind-code the predictor and report the sighted-versus-blind gap; separate mechanical inputs from outcome-encoding ones and report the conservative index; enter conditioning covariates as controls rather than filters; and code the load-bearing cases against the hypothesis where the record is contested. A claim that survives all four has earned attention. The circulation criterion did not.

Three limits bound the conclusion, and overstating a null is as much an error as overstating a positive. The result is for one operationalisation of circulation-centrality; a different and genuinely non-circular construct could be built and might fare better, though the burden now sits with its proposer. The sample is small, and a larger panel of rescue decisions, including the many quiet failures that never provoked a salient decision, could revisit the question with more power. The capacity and fiscal-mandate covariates have been coded both by the author and, independently, by a coder blind to the hypothesis, and the null is robust to the independent coding (Table 1); the one remaining refinement is independent authorship of the entity profiles, which the blind index does not yet have. What the paper establishes is narrower than “circulation never matters”. It is that a specific, carefully built, blind-coded test of the circulation criterion does not out-predict ordinary systemic importance on this panel, and that the apparent evidence to the contrary was an artifact whose construction can be shown step by step.

8. Conclusion

The boundary of the safety net invites a clean theory, and the cleanest on offer was that a central bank rescues what is central to monetary circulation rather than what is merely large or legally favoured. This paper built that theory into a falsifiable test and watched it fail. A point-biserial correlation of 0.63 fell to 0.13 as three disciplines were applied: blind coding removed the author’s hindsight, a plumbing-only index removed two components that restated the outcome, and the documentary record on Lehman and Washington Mutual removed the rest. What predicts rescue, on these thirty decisions, is conventional systemic importance and the political-legal capacity to act, the too-big-to-fail account the test set out to beat.

A negative result of this kind is not a dead end. It tells us that the boundary of last-resort lending is not hiding in the plumbing, that the next theory of it must beat size and politics on cases coded against itself, and that strong-looking findings in this domain should be met with the three questions that undid this one: is the predictor blind to the outcome, are the inconvenient cases controlled for rather than deleted, and do the load-bearing cases survive the record? The symbiosis metaphor that motivated the test remains a vivid way to describe the dependence of private money on public; what it does not appear to be is a source of predictions that ordinary systemic importance does not already supply.


References

Acharya, V. V., and T. Yorulmazer. 2007. “Too Many to Fail: An Analysis of Time-Inconsistency in Bank Closure Policies.” Journal of Financial Intermediation 16 (1).

Aglietta, M., and A. Orléan. 1998. La Monnaie souveraine. Paris: Odile Jacob.

Ball, L. M. 2018. The Fed and Lehman Brothers: Setting the Record Straight on a Financial Disaster. Cambridge: Cambridge University Press. (NBER Working Paper 22410, 2016.)

Buiter, W., et al. 2023. “Stabilising Financial Markets: Lending and Market Making as a Last Resort.” SSRN Electronic Journal.

Financial Crisis Inquiry Commission (FCIC). 2011. The Financial Crisis Inquiry Report. Washington, DC: US Government Printing Office.

Kane, E. J. 1981. “Accelerating Inflation, Technological Innovation, and the Decreasing Effectiveness of Banking Regulation.” Journal of Finance.

Mehrling, P. 2011. The New Lombard Street: How the Fed Became the Dealer of Last Resort. Princeton: Princeton University Press.

Nersisyan, Y. 2015. “The Repeal of the Glass–Steagall Act and the Federal Reserve’s Extraordinary Intervention during the Global Financial Crisis.” Journal of Post Keynesian Economics.

Schwarcz, S. L. 2010. “Too Big to Fail? Recasting the Financial Safety Net.” SSRN Electronic Journal.

Searle, J. R. 1995. The Construction of Social Reality. New York: Free Press.