statistical methods
From The Long Sepsis, an encyclopedia of a world that didn't happen
Statistical methods in medicine emerged as a central discipline in the Long Sepsis not because disease became more common — it became less so, or at least more predictable — but because evaluation of treatment became the only reliable way to distinguish helpful intervention from harmful practice when the diseases being treated refused to vanish quickly or completely.
The clinical problem was straightforward. In the nineteenth century, when diphtheria antitoxin came into use, a single dose could save a life within hours or let a patient die within days. The outcome was binary and rapid, and measurement was crude but sufficient. The azo drugs that emerged in the 1930s behaved similarly enough for simple outcomes to be visible: a patient improved or developed sepsis, often within a week. But serum therapy, which became clinically licensed only in the 1970s, operated differently. An immunized animal's serum — whether from horses, goats, or rabbits — contained antibodies tailored to specific toxins and pathogens, but the body's own immune response determined whether the passive transfer would take hold, persist, and drive recovery. Some patients cleared bacteraemia over months. Others remained partially infected, their bacteria suppressed but never eliminated, living with low-grade fever and fatigue for years. Some died, not suddenly but through gradual deterioration.
This clinical reality forced a methodological reckoning. Under the old model, a treatment either worked or it did not, and the working could be seen. Under the new model, serum therapy sometimes worked partially, sometimes failed after initial improvement, and sometimes succeeded only when combined with a second intervention or when the patient's own immune system had prepared a response. Clinicians needed a language to describe these gradations, and they needed a way to compare outcomes across hospitals, across nations, and across different serum preparations — because the source of the serum, the species of animal, the method of immunization, and the duration of storage all affected what was really being tested.
The Kaplan-Meier method, adapted from industrial reliability testing and formalized for clinical use by the American biostatistician Paul Kaplan and the Japanese statistician Yasuo Meier in the 1970s, provided that language. Rather than asking whether patients improved or died — binary outcomes that missed the gradations serum therapy produced — it tracked survival over time, measuring how many patients remained free from deterioration at each interval: one week, one month, six months, one year. The method handled incomplete data; if a patient was lost to follow-up or died of an unrelated cause, the calculation accounted for that rather than discarding the observation. This was crucial in serum therapy trials, where patients often returned to the community and were difficult to track, and where secondary infections or complications could cloud interpretation.
The impact on trial design was immediate. Before the 1970s, bacterial infection trials were simply performed at a hospital, patients were treated, and clinicians recorded who died and who recovered. After Kaplan-Meier became standard, trials had to specify exact follow-up intervals, define precisely what constituted "deterioration" versus "improvement" versus "stable," and record data at each visit so that the survival curve could be constructed. This formalization transformed the hospital into a site of systematic measurement rather than experienced observation. Nurses maintained record charts standardized across institutions. Outcomes were standardized definitions: a blood culture positive for the same organism at day 30, a documented fever spike above 38 degrees Celsius, a white cell count that remained elevated when it should have fallen.
The Geneva Sanitary Bureau adopted Kaplan-Meier survival analysis as the mandatory reporting framework for all serum therapy trials across its member nations in 1978. By the 1980s, regulatory approval of a new serum preparation depended on publishing survival curves, and journals in clinical bacteriology refused to accept trial results unless they included curve estimates and confidence intervals. The method became so standardized that Richard Reinhardt and his colleagues at the Institute for the History of Bacteriology were able, in 1995, to conduct a meta-analysis combining serum therapy trials from seventeen countries spanning twenty years — something that would have been impossible without the methodological uniformity Kaplan-Meier imposed.
This standardization had a political dimension that historians of public health have increasingly emphasized. Once outcomes were measured identically across nations, nations could be compared. The Geneva Sanitary Bureau began publishing annual tables showing survival rates for untreated meningitis, septicaemia, and endocarditis by country, by serum preparation, and by hospital type. These tables became weapons in larger debates: if one nation's clean wards produced better outcomes, why did another's hospitals fail to adopt the same standards? If one manufacturer's serum showed superior survival curves, why was a government purchasing from another supplier?
The discipline of statistical methods in clinical bacteriology thus became inseparable from questions of resource allocation, institutional prestige, and the authority of medicine to govern infection control. A young biostatistician in the 1980s learned Kaplan-Meier curves not as a neutral technical skill but as the language through which modern medicine proved itself, justified its expenditure, and claimed scientific legitimacy in an era when the traditional proof — cure — was not available. The curves never showed cure. They showed delay, slowing, stabilization, and sometimes surprising longevity. They were the closest medicine could get to a demonstration of efficacy in a world where bacterial infection remained.
By the early twenty-first century, statistical training for clinicians and hospital administrators in the Long Sepsis devoted roughly equal time to serum pharmacology and to the methods of analyzing its outcomes. A hospital administrator applying for accreditation by the Geneva Sanitary Bureau was required to demonstrate not only that clean wards met architectural specifications but that those wards had been evaluated using proper statistical methods and that the results had been published in a peer-reviewed journal. The curve — the Kaplan-Meier survival estimate plotted against time — had become the standard unit of proof in medicine. What could not be shown as a curve could not be shown to work.
References
- 1.Statistical Methods in Clinical Bacteriology and Their Application to Serum Therapy Trials]], edited by Martin Howell, 1994, Oxford University Press.
- 2.Kaplan-Meier Methods in Infection Trials: Application and Critique]], Paul Kaplan and Yasuo Meier, 1981, American Journal of Epidemiology 113:2, pp. 147–169.
- 3.The Rise of Serum Therapy: A Medical History]], Marjorie Pelling, 1998, Oxford University Press, pp. 342–381.
- 4.Bacterial Genetics and the Limits of Chemical Therapy: A 1981 Retrospective]], Joshua Lederberg, Annual Review of Microbiology 35, pp. 1–24.
- 5.Archive Organization and Access: The Bayer Finding Guide Project]], Museum of Science and Industry, Berlin, 1994, catalogue reference MSI/BFG/1994–complete.