tableone: An open source Python package for producing summary statistics for research papers

tableone: An open source Python package for producing summary statistics for research papers
复制标题

DOI:
10.1093/jamiaopen/ooy012
复制
发表时间:
2018-07-01
期刊:
影响因子:
2.1
通讯作者:
Mark, Roger G.
Mark, Roger G.
中科院分区:
其他
文献类型:
--
作者:
Pollard, Tom J.;Johnson, Alistair E. W.;Mark, Roger G.

文献摘要

被引文献

相似文献

目标:在定量研究中,了解研究人群的基本参数是解释结果的关键。因此,研究论文的第一个表(“表 1”)通常包含研究数据的汇总统计数据。我们的目标有两个。首先,我们寻求提供一种简单、可重复的方法来提供 Python 编程语言研究论文的汇总统计数据。其次,我们寻求使用该软件包来提高研究论文中报告的汇总统计数据的质量。材料和方法:tableone 软件包是按照科学计算的良好实践指南开发的,所有代码均在 MIT 许可下提供。测试框架运行在持续集成服务器上,有助于保持代码稳定性。公开跟踪问题并鼓励公众贡献。结果:tableone 软件包自动将汇总统计数据编译为可发布的格式,例如 CSV、HTML 和 LaTeX。可执行的 Jupyter Notebook 演示了该包对 MIMIC-III 数据库中的数据子集的应用。计算诸如用于异常值检测的 Tukey 规则和用于模态的 Hartigan's Dip 测试等测试,以突出总结数据时的潜在问题。 讨论和结论:我们为研究人员提供开源软件,以方便使用 Python(科学研究中日益流行的语言)进行可重复的研究。该工具包旨在通过社区反馈和输入随着时间的推移而成熟。开发用于汇总数据的通用工具作为现有指南和建议的补充可能有助于促进良好实践。我们鼓励将 tableone 与其他描述性统计方法一起使用,特别是可视化,以确保适当的数据处理。我们还建议在使用 tableone 进行研究时寻求统计学家的指导,尤其是在提交研究发表之前。
Objectives: In quantitative research, understanding basic parameters of the study population is key for interpretation of the results. As a result, it is typical for the first table ("Table 1") of a research paper to include summary statistics for the study data. Our objectives are 2-fold. First, we seek to provide a simple, reproducible method for providing summary statistics for research papers in the Python programming language. Second, we seek to use the package to improve the quality of summary statistics reported in research papers.Materials and Methods: The tableone package is developed following good practice guidelines for scientific computing and all code is made available under a permissive MIT License. A testing framework runs on a continuous integration server, helping to maintain code stability. Issues are tracked openly and public contributions are encouraged.Results: The tableone software package automatically compiles summary statistics into publishable formats such as CSV, HTML, and LaTeX. An executable Jupyter Notebook demonstrates application of the package to a subset of data from the MIMIC-III database. Tests such as Tukey's rule for outlier detection and Hartigan's Dip Test for modality are computed to highlight potential issues in summarizing the data.Discussion and Conclusion: We present open source software for researchers to facilitate carrying out reproducible studies in Python, an increasingly popular language in scientific research. The toolkit is intended to mature over time with community feedback and input. Development of a common tool for summarizing data may help to promote good practice when used as a supplement to existing guidelines and recommendations. We encourage use of tableone alongside other methods of descriptive statistics and, in particular, visualization to ensure appropriate data handling. We also suggest seeking guidance from a statistician when using tableone for a research study, especially prior to submitting the study for publication.