Toward Training and Assessing Reproducible Data Analysis in Data Science Education
摘要 : Reproducibility is a cornerstone of scientific research. Data science is not an exception. In recent years scientists were concerned about a large number of irreproducible studies. Such reproducibility crisis in science could severely undermine public trust in science and science-based public policy. Recent efforts to promote reproducible research mainly focused on matured scientists and much less on student training. In this study, we conducted action research on students in data science to evaluate to what extent students are ready for communicating reproducible data analysis. The results show that although two-thirds of the students claimed they were able to reproduce results in peer reports, only one-third of reports provided all necessary information for replication. The actual replication results also include conflicting claims; some lacked comparisons of original and replication results, indicating that some students did not share a consistent understanding of what reproducibility means and how to report replication results. The findings suggest that more training is needed to help data science students communicating reproducible data analysis.
[V1] | 2022-11-29 13:31:02 | ChinaXiv:202211.00463V1 | 下载全文 |
1. 人工智能在匿名网络追踪中网站指纹的应用综述 2023-05-30 |
2. Copula熵:理论和应用 2023-05-19 |
3. A Preliminary Study on the Capability Boundary of LLM and a New Implementation Approach for AGI 2023-05-06 |
4. DASICS-安全处理器设计白皮书 2023-04-18 |
5. 对人工智能大模型能力边界的初探和一种新的AGI实现途径 2023-04-18 |
6. DASICS-安全处理器设计白皮书 2023-04-18 |
7. DASICS-安全处理器设计白皮书 2023-04-17 |
8. Call for urgent regulations on artificial intelligence 2023-04-15 |
9. An Improved YOLOv5-Based Method for UAV Object Detection 2023-03-23 |
10. Curvature-Balanced Feature Manifold Learning for Long-Tailed Classification 2023-03-22 |