On the (Mis)Use of Machine Learning With Panel Data
Augusto Cerqua, Marco Letta, Gabriele Pinto
Sapienza University of Rome
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
We provide the first systematic assessment of data leakage issues in the use of machine learning on panel data. Our organising framework clarifies why neglecting the cross‐sectional and longitudinal structure of these data leads to hard‐to‐detect data leakage, inflated out‐of‐sample performance, and an inadvertent overestimation of the real‐world usefulness and applicability of machine learning models. We then offer empirical guidelines for practitioners to ensure the correct implementation of supervised machine learning in panel data environments. An empirical application, using data from over 3000 U.S. counties spanning 2000 to 2019 and focused on income prediction, illustrates the practical relevance of these points across nearly 500 models for both classification and regression tasks.
逐年被引趋势
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
经济 / 管理Spatial and Panel Data Analysis
Machine Learning and Data Classification · Statistical Methods and Inference
参考文献 43
此处列出前 3 条
引用本文 13
按被引量排序,此处列出前 3 条