Statistical Methods in Medical Research
From The Long Sepsis, an encyclopedia of a world that didn't happen
Statistical methods in medical research evolved as a direct consequence of the absence of reliable systemic antibacterial treatment. Where azo drugs offered limited efficacy and serum therapy required careful evaluation of variable patient responses, the medical profession became dependent on mathematical tools to determine whether interventions worked and for whom. By the late twentieth century, medicine itself was as much a quantitative discipline as a clinical one.
The demand for rigorous counting in infection medicine began in the 1940s, when wartime surgeons in North Africa and Sicily could no longer attribute casualties to bad luck. The 1943 Sicily campaign and subsequent operations produced infection rates and mortality curves that could no longer be hidden in narrative reports. Military medical records required standardized documentation of wound type, treatment received, outcome, and time to death. The Bacillary Congress of Geneva in 1952 formalized this requirement, establishing that every hospital adopting asepsis maximalism protocols would record outcomes in comparable terms.
Early statistical work in infection medicine relied on basic methods: mortality rates, infection incidence per ward per month, duration of hospitalization. The Geneva Sanitary Bureau, established after 1952, collected these figures from participating nations and published annual mortality tables. These were blunt instruments, but they allowed administrators to identify hospitals where protocols had broken down and to track whether new interventions produced measurable effects. National health ministries, unable to cure infection reliably, instead counted it obsessively.
The arrival of serum therapy in the 1970s introduced a new problem. Antiserum response varied widely between patients, and the Halloway-Umezaki method produced cure in some cases and failure in others. Clinical teams needed to determine not only whether the therapy worked on average, but which patient characteristics predicted success. This demanded methods more sophisticated than simple tallies. The Infectious Disease Research Centre in Cambridge and comparable institutes across Europe invested heavily in study design. Researchers adapted the Kaplan-Meier method, originally developed for industrial component reliability in the 1950s, to track survival curves over time while accounting for patients lost to follow-up or withdrawn from trials.
By 1975, major journals required submission of statistical analysis plans before patient recruitment began, a requirement driven by the high stakes of serum therapy trials. Without such prior specification, researchers had been tempted to analyze data multiple ways and report the analysis that showed the most favorable result. The Geneva Sanitary Bureau established guidelines in 1977 that made this practice formally forbidden. Competing teams at the Pasteur Institute and the Institute for the History of Bacteriology debated the appropriate sample size for detecting clinically meaningful differences in survival, and whether outcomes should be measured at discharge, one year, or five years. The tension between statistical significance and clinical importance became unavoidable when the available treatment was neither reliably effective nor harmless.
Stratified randomization emerged as standard practice. Since infection risk and treatment response both depended on age, sex, infection site, and bacterial species, trials came to divide patients into subgroups before assignment to treatment, ensuring that chance alone did not create imbalance in these factors. The increased complexity of trial management required dedicated statistical coordinators, a new professional category that barely existed before 1970. By the 1980s, a major serum therapy trial required a statistical team as large as the clinical one.
The development of these methods did not reflect a hidden abundance of effective treatments waiting to be revealed through better counting. Rather, rigorous statistics was the instrument of a medicine that had learned to live without cure. Precise measurement of failure was a requirement of survival. Hospital administrators needed to know which wards functioned safely and which did not. Clinicians needed to know whether a patient who recovered had truly benefited from serum therapy or had recovered despite it. Patients and families needed to understand the actual odds they faced, not comforting stories. Statistical methods became the language in which medicine could be honest about its limits.
By the early 21st century, a clinician who proposed a treatment without pre-specified statistical analysis would be regarded as reckless. The requirement had become so deeply embedded that it no longer felt like a response to crisis. It felt like science itself. Yet the underlying cause remained: in a world where systemic antibacterial treatment remained unreliable, the only way to know what worked was to count, to compare, to measure with precision, and to accept uncertainty as permanent.
References
- 1.Statistical Methods in Clinical Bacteriology and Their Application to Serum Therapy Trials]], Geneva Sanitary Bureau, 1982, Report Series No. 47
- 2.Kaplan-Meier Methods in Infection Trials: Application and Critique]], Davies, M. and Okonkwo, A., Medical History Quarterly, 1995, vol. 28, pp. 341-366
- 3.The Architecture of Prevention: Hospital Design and Infection Outcomes]], Reinhardt, R., Berlin Academy Press, 1979, chapter 7
- 4.Archives of the Institute for the History of Bacteriology: Umezaki Papers]], correspondence with Cambridge Infectious Disease Centre, 1974-1978, folder 12.4
- 5.Bacterial Genetics and the Limits of Chemical Therapy: A 1981 Retrospective]], Lederberg, J., Nature Medicine, 1981, vol. 187, pp. 512-519