142
25
204
34
报告摘要 Offline policy learning aims at utilizing observations collected a priori (from either fixed or adaptively evolving behavior policies) to learn an optimal individualized decision rule that achieves the best overall outcomes for a given population. Existing policy learning methods rely on a uniform overlap assumption, i.e., the propensities of exploring all actions for all individual characteristics must be lower bounded. As one has no control over the data collection process, this assumption can be unrealistic in many situations, especially when the behavior policies are allowed to evolve over time with diminishing propensities for certain actions.  In this work, we propose Pessimistic Policy Learning (PPL), a new algorithm that optimizes lower confidence bounds (LCBs) -- instead of point estimates -- of the policy values. In our theoretical analysis, we develop a new self-normalized type concentration inequality for inverse-propensity-weighting estimators, generalizing the well-known empirical Bernstein's inequality to unbounded and non-i.i.d.~data. We complement our theory with an efficient optimization algorithm via Majorization-Minimization and policy tree search, as well as extensive experiments that demonstrate the efficacy of PPL. 嘉宾简介 Ying Jin is a fifth-year PhD candidate in the Department of Statistics at Stanford University, advised by Professors Emmanuel Candès and Dominik Rothenhäusler. Her research interests include conformal prediction, selective inference, distribution robustness, and data-driven decision-making. 直播分享时间:2024年4月27日
统计学第二课堂
狗熊会
214/374
loading
加载中...
上海财经大学张耀武教授:高维数据中的非线性关系和独立性检验
2918播放
中国人民大学张琨助理教授:条件风险值的模拟置信区间
1609播放
中国人民大学在读博士闫引桥:空间转录组学研究中的贝叶斯整合区域分割方法
2400播放
ChatGPT辅助的R语言编程:06-生存回归
1696播放
playing
斯坦福大学在读博士生金滢:无重叠的政策学习-悲观主义和广义经验伯恩斯坦不等式
9117播放
香港中文大学范青亮副教授:带有多个无效及弱工具变量的内生性处理效应模型
1985播放
西南财经大学刘耀午教授:全局检验的集成方法
1799播放
中央财经大学李丰副教授:基于狄利克雷过程的无限预测组合
2190播放
复旦大学在读博士任怡萌:大规模网络下空间自回归模型的分布式估计与推断方法
3.8万播放
上海科技大学汪时嘉助理教授:复杂模型的近似贝叶斯加速计算方法
3232播放
ChatGPT辅助的R语言编程:05-泊松回归
1694播放
中国人民大学王霞教授:区分时变因子模型
5834播放
ChatGPT辅助的R语言编程:04-定序回归
1508播放
【小丫聊数据】数据会说谎?统计学教授一分钟讲清楚辛普森悖论
2061播放
第三届美团商业分析精英大赛季军作品:“袋鼠管家”新模式——基于遗传算法的无人配送车最优投放方案
1806播放
第三届美团商业分析精英大赛季军作品:腰部博主的广告投放价值研究——基于博主的历史动态和视频信息数据
1834播放
第三届美团商业分析精英大赛季军作品:基于强化学习的新能源充电站布局优化
1712播放
第三届美团商业分析精英大赛亚军作品:助力“闪电仓”老品去库存——临期食品动态定价与管理策略
2094播放
第三届美团商业分析精英大赛亚军作品:基于数据驱动的直播卖券策略与产品设计优化-瑞幸咖啡在抖音直播电商平台的创新拓展与市场增长分析
1940播放
第三届美团商业分析精英大赛冠军作品:Meta-tour旅游定制师——整合多源信息的一站式旅游规划研究
2961播放
第二届美团商业分析精英大赛季军作品:外卖“飞上天”——基于成本和需求优化的无人机起飞机场选址研究
1806播放
第二届美团商业分析精英大赛季军作品:智慧零售时代-基于需求预测的无人微仓联合补货管理系统
1827播放
第二届美团商业分析精英大赛季军作品:基于历史销量数据的 “云厨房”业务模式可行性探析
1534播放
第二届美团商业分析精英大赛亚军作品:高端现制茶饮品牌的选址问题——基于市场下沉背景
1634播放
loading
加载中...
客服
顶部