ChemRxiv
These are preliminary reports that have not been peer-reviewed. They should not be regarded as conclusive, guide clinical practice/health-related behavior, or be reported in news media as established information. For more information, please see our FAQs.
manuscript.pdf (8.25 MB)

Powerful Statistical Tests for Ordered Data

preprint
revised on 20.04.2021, 15:52 and posted on 21.04.2021, 05:22 by Juergen Koefinger, Gerhard Hummer

The inference of models from one-dimensional ordered data subject to noise is a fundamental and ubiquitous task in the physical and life sciences. A prototypical example is the analysis of small- and wide-angle solution scattering experiments using x-rays (SAXS/WAXS) or neutrons (SANS). In such cases, it is common practice to check the quality of a fit by using Pearson's chi-square test, which ignores the order of the data. We usually plot the residuals and check visually for systematic deviations without quantifying them. To quantify these deviations, we developed test statistics based on the distributions of the lengths of the runs of the signs of the residuals. Specifically, we use the probability of run-length distributions, for which we provide analytical expressions, to rank them and to calculate their P-values. We introduce the Shannon information distribution as an elegant and versatile tool for calculating P-values. We find that these distributions follow shifted gamma distributions, such that they are summarized by three parameters only. We show for a set of six models that our test statistics are more powerful than Pearson's chi-square test and common sign-based tests. We provide an open source Python 3 implementation of our tests free of charge at https://github.com/bio-phys/hplusminus.

Funding

Max Planck Society

History

Email Address of Submitting Author

juergen.koefinger@biophys.mpg.de

Institution

Max Planck Insitute of Biophysics

Country

Germany

ORCID For Submitting Author

0000-0001-8367-1077

Declaration of Conflict of Interest

No conflict of interest.

Version Notes

Updated with SASBDB analysis. Improved writing.

Exports