- Acquired (You could look at its transcript using iPhone.)
7/13/2026
一些常聽的 Podcast 節目和培養英文聽力的方法
7/13/2025
Statistical Modeling: The Two Cultures
Cynthia Rudin, Leo Breiman, the Rashomon Effect, and the Occam Dilemma, arXiv:2507.03884, 2025.
In the famous “Two Cultures” paper, Leo Breiman provided a visionary perspective on the cultures of “data models” (modeling with consideration of data generation) versus “algorithmic models” (vanilla machine learning models). I provide a modern perspective on these two approaches. One of Breiman’s key arguments against data models is what he called the “Rashomon Effect,” which is the existence of many different-but-equally-good models. The Rashomon Effect implies that data modelers would not be able to determine which model generated the data. Conversely, one of his core advantages in favor of data models is simplicity, as he claimed there exists an “Occam Dilemma,” i.e., an accuracy-simplicity tradeoff, where algorithmic models must be complex in order to be accurate. After 25 years of more powerful computers, it has become clear that this claim is not generally true, in that algorithmic models do not need to be complex to be accurate; however, there are nuances that help explain Breiman’s logic, specifically, that by “simple,” he appears to consider only linear models or unoptimized decision trees. Interestingly, the Rashomon Effect is a key tool in proving the nullification of the Occam Dilemma. To his credit though, Breiman did not have the benefit of modern computers, with which my observations are much easier to make.
4/13/2025
張忠謀自傳 (下冊)
張忠謀,《張忠謀自傳》下冊 (1964-2018),天下,2024
「下冊」的第一篇「德儀篇」占全書三分之一篇幅,敘述我自 27 歲至 52 歲黃金年齡的崛起及衰退。德儀時期也是我最為熱情澎湃的年代。德儀讓我「獨上高樓,望盡天涯路」;為此,我無比地感激。但是,它對我的培植也激發了我要登上德儀顛峰的雄心。但是,這雄心是無法實現的;不說別的,只看 1970 年代美國德克薩斯州的客觀環境,就可斷定我的雄心很難很難實現。但是我的痴心,居然使我「明知不可為」的情況下,還在德儀多滯留了 5、6 年,後來連婚姻都賠上了。現在回想,那 5、6 年是多麼絕望的 5、6 年!
1/26/2025
價值戰爭
作者:朱敬一/ 羅昌發/ 李柏青/ 林建志,譯者:許瑞宋,價值戰爭: 極權中國與民主陣營的終極經濟衝突,衛城出版,2023
美中經濟衝突的根源是什麼?
經貿往來為什麼無法造就民主國家與專制國家的穩定合作?
何以市場理性無法擱置自由與集權的對峙?
10/16/2024
Daron Acemoglu is not having all this AI hype
Robin Wigglesworth, Daron Acemoglu is not having all this AI hype, Financial Times, May 28 2024.
Who is Daron Acemoglu? 246,645 Citations so far!
2/26/2024
Catastrophe Insurance Pricing
C. Zeng and D. Bertsimas, Catastrophe Insurance Pricing: A Robust Optimization Approach, In preparation for Management Science, 2023.
2/25/2024
Customer Choice Models vs. Machine Learning
Jacob Feldman, Dennis J. Zhang, Xiaofei Liu, and Nannan Zhang (2021) Customer Choice Models vs. Machine Learning: Finding Optimal Product Displays on Alibaba. Operations Research 70(1):309-328. (Best OM Paper in Operations Research Award: Finalist, pdf, implementation details)
2/23/2024
2023 Franz Edelman Award
2023 Edelman Competition (video)
Prakhar Mehrotra et al., (2024) Optimizing Walmart’s Supply Chain from Strategy to Execution. INFORMS Journal on Applied Analytics 54(1):5-19. (2023 Franz Edelman Award) (Keywords: supply chain optimization, network design, simulation, truck routing and loading, mixed-integer programming, metaheuristics)
2/18/2024
An Exact Solution to Wordle
Dimitris Bertsimas, Alex Paskov (2024) An Exact Solution to Wordle. Operations Research.
1/19/2024
How to Avoid a Climate Disaster (如何避免氣候災難)
Bill Gates, How to Avoid a Climate Disaster: The Solutions We Have and the Breakthroughs We Need, Random House, February 23, 2021.
1/09/2024
HBR at 100
Harvard Business Review, HBR at 100: The Most Influential and Innovative Articles from Harvard Business Review's First Century, June 14, 2022.
哈佛商業評論,哈佛商業評論最有影響力的30篇文章,天下文化,2022/6/30
7/22/2023
The role of optimization in some recent advances in data-driven decision-making
Baardman, L., Cristian, R., Perakis, G. et al. The role of optimization in some recent advances in data-driven decision-making. Mathematical Programming 200, 1–35 (2023). https://doi.org/10.1007/s10107-022-01874-9.
Data-driven decision-making has garnered growing interest as a result of the increasing availability of data in recent years. With that growth many opportunities and challenges have sprung up in the areas of predictive and prescriptive analytics. Often, optimization can play an important role in tackling these issues. In this paper, we review some recent advances that highlight the difference that optimization can make in data-driven decision-making. We discuss some of our contributions that aim to advance both predictive and prescriptive models. First, we describe how we can optimally estimate clustered models that result in improved predictions. Next, we consider how we can optimize over objective functions that arise from tree ensemble models in order to obtain better prescriptions. Finally, we discuss how we can learn optimal solutions directly from the data allowing for prescriptions without the need for predictions. For all these new methods, we stress the need for good performance but also the scalability to large heterogeneous datasets.
4/17/2023
A Practical End-to-End Inventory Management Model with Deep Learning
Meng Qi, Yuanyuan Shi, Yongzhi Qi, Chenxin Ma, Rong Yuan, Di Wu, Zuo-Jun (Max) Shen (2023) A Practical End-to-End Inventory Management Model with Deep Learning. Management Science 69(2):759-773. (Data and Python codes)
We investigate a data-driven multiperiod inventory replenishment problem with uncertain demand and vendor lead time (VLT) with accessibility to a large quantity of historical data. Different from the traditional two-step predict-then-optimize (PTO) solution framework, we propose a one-step end-to-end (E2E) framework that uses deep learning models to output the suggested replenishment amount directly from input features without any intermediate step. The E2E model is trained to capture the behavior of the optimal dynamic programming solution under historical observations without any prior assumptions on the distributions of the demand and the VLT. By conducting a series of thorough numerical experiments using real data from one of the leading e-commerce companies, we demonstrate the advantages of the proposed E2E model over conventional PTO frameworks. We also conduct a field experiment with JD.com, and the results show that our new algorithm reduces holding cost, stockout cost, total inventory cost, and turnover rate substantially compared with JD’s current practice. For the supply chain management industry, our E2E model shortens the decision process and provides an automatic inventory management solution with the possibility to generalize and scale. The concept of E2E, which uses the input information directly for the ultimate goal, can also be useful in practice for other supply chain management circumstances.
3/15/2023
2022 Franz Edelman Award
2022 Edelman Competition (videos)
Leonardo J. Basso et al., Analytics Saves Lives During the COVID-19 Crisis in Chile, INFORMS Journal on Applied Analytics, 2023, 53(1):9-31. (2022 Franz Edelman Award) (statistical analysis, integer programming, regression)
During the COVID-19 crisis, the Chilean Ministry of Health and the Ministry of Sciences, Technology, Knowledge and Innovation partnered with the Instituto Sistemas Complejos de Ingeniería (ISCI) and the telecommunications company ENTEL, to develop innovative methodologies and tools that placed operations research (OR) and analytics at the forefront of the battle against the pandemic. These innovations have been used in key decision aspects that helped shape a comprehensive strategy against the virus, including tools that (1) provided data on the actual effects of lockdowns in different municipalities and over time; (2) helped allocate limited intensive care unit (ICU) capacity; (3) significantly increased the testing capacity and provided on-the-ground strategies for active screening of asymptomatic cases; and (4) implemented a nationwide serology surveillance program that significantly influenced Chile’s decisions regarding vaccine booster doses and that also provided information of global relevance. Significant challenges during the execution of the project included the coordination of large teams of engineers, data scientists, and healthcare professionals in the field; the effective communication of information to the population; and the handling and use of sensitive data. The initiatives generated significant press coverage and, by providing scientific evidence supporting the decision making behind the Chilean strategy to address the pandemic, they helped provide transparency and objectivity to decision makers and the general population. According to highly conservative estimates, the number of lives saved by all the initiatives combined is close to 3,000, equivalent to more than 5% of the total death toll in Chile associated with the pandemic until January 2022. The saved resources associated with testing, ICU beds, and working days amount to more than 300 million USD.
3/13/2023
Exploring the Whole Rashomon Set of Sparse Decision Trees
Rui Xin, Chudi Zhong, Zhi Chen, Takuya Takagi, Margo Seltzer, Cynthia Rudin, Exploring the Whole Rashomon Set of Sparse Decision Trees, NeurIPS (oral), 2022. (code) | (bib) | (5 min video)
In any given machine learning problem, there may be many models that could explain the data almost equally well. However, most learning algorithms return only one of these models, leaving practitioners with no practical way to explore alternative models that might have desirable properties beyond what could be expressed within a loss function. The Rashomon set is the set of these all almost-optimal models. Rashomon sets can be extremely complicated, particularly for highly nonlinear function classes that allow complex interaction terms, such as decision trees. We provide the first technique for completely enumerating the Rashomon set for sparse decision trees; in fact, our work provides the first complete enumeration of any Rashomon set for a non-trivial problem with a highly nonlinear discrete function class. This allows the user an unprecedented level of control over model choice among all models that are approximately equally good. We represent the Rashomon set in a specialized data structure that supports efficient querying and sampling. We show three applications of the Rashomon set: 1) it can be used to study variable importance for the set of almost-optimal trees (as opposed to a single tree), 2) the Rashomon set for accuracy enables enumeration of the Rashomon sets for balanced accuracy and F1-score, and 3) the Rashomon set for a full dataset can be used to produce Rashomon sets constructed with only subsets of the data set. Thus, we are able to examine Rashomon sets across problems with a new lens, enabling users to choose models rather than be at the mercy of an algorithm that produces only a single model.
2/19/2023
思維的製程
彭建文,思維的製程:台積電教我的思維進階法,練成全局經營腦和先進工作術,商業周刊,2023
職場不怕碰上難題,怕的是不會聰明解決。
學習台積電多年淬鍊的系統性問題解決策略,
你也能優化思維的製程,扎穩職涯腳步,累積世界第一競爭力。
2/01/2023
Generalized Synthetic Control for TestOps at ABI
Luis Costa, Vivek F. Farias, Patricio Foncea, Jingyuan (Donna) Gan, Ayush Garg, Ivo Rosa Montenegro, Kumarjit Pathak, Tianyi Peng, and Dusan Popovic, Generalized Synthetic Control for TestOps at ABI: Models, Algorithms, and Infrastructure, To appear in INFORMS Journal on Applied Analytics (Winner, Daniel H. Wagner Prize 2022)
We describe a novel optimization-based approach– Generalized Synthetic Control (GSC)– to learning from experiments conducted in the world of physical retail. GSC solves a long-standing problem of learning from physical retail experiments when treatment effects are small, the environment is highly noisy and non-stationary, and interference and adherence problems are commonplace. The use of GSC has been shown to yield an approximately 100x increase in power relative to typical inferential methods and forms the basis of a new large-scale testing platform: ‘TestOps’. TestOps was developed and has been broadly implemented as part of a collaboration between Anheuser Busch Inbev (ABI) and an MIT team of operations researchers and data engineers. TestOps currently runs physical experiments impacting approximately 135M USD in revenue every month and routinely identifies innovations that result in a 1-2% increase in sales volume. The vast majority of these innovations would have remained unidentified absent our novel approach to inference: prior to our implementation, statistically significant conclusions could be drawn on only ∼ 6% of all experiments; a fraction that has now increased by over an order of magnitude.
12/09/2022
為什麼聰明人會做蠢事 ?
姚怡平譯,為什麼聰明人會做蠢事? 顛覆高智商等於絕對聰明的常理,助你找出決策的關鍵智慧,商業周刊,2020
David Robson, The Intelligence Trap: Revolutionise Your Thinking and Make Wiser Decisions, Hodder & Stoughton, 2019.
Waki 瓦基,《為什麼聰明人會做蠢事?》讀書心得:掌握變聰明的3種方法,2021-09-01 (也有Podcast)
11/01/2022
10/27/2022
Stroke risk is not linear
Orfanoudaki A, Chesley E, Cadisch C, Stein B, Nouh A, Alberts MJ, et al. (2020) Machine learning provides evidence that stroke risk is not linear: The non-linear Framingham stroke risk score. PLoS ONE 15(5): e0232414. https://doi.org/10.1371/journal.pone.0232414
Current stroke risk assessment tools presume the impact of risk factors is linear and cumulative. However, both novel risk factors and their interplay influencing stroke incidence are difficult to reveal using traditional additive models. The goal of this study was to improve upon the established Revised Framingham Stroke Risk Score and design an interactive Non-Linear Stroke Risk Score. Leveraging machine learning algorithms, our work aimed at increasing the accuracy of event prediction and uncovering new relationships in an interpretable fashion. A two-phase approach was used to create our stroke risk prediction score. First, clinical examinations of the Framingham offspring cohort were utilized as the training dataset for the predictive model. Optimal Classification Trees were used to develop a tree-based model to predict 10-year risk of stroke. Unlike classical methods, this algorithm adaptively changes the splits on the independent variables, introducing non-linear interactions among them. Second, the model was validated with a multi-ethnicity cohort from the Boston Medical Center. Our stroke risk score suggests a key dichotomy between patients with history of cardiovascular disease and the rest of the population. While it agrees with known findings, it also identified 23 unique stroke risk profiles and highlighted new non-linear relationships; such as the role of T-wave abnormality on electrocardiography and hematocrit levels in a patient’s risk profile. Our results suggested that the non-linear approach significantly improves upon the baseline in the c-statistic (training 87.43% (CI 0.85–0.90) vs. 73.74% (CI 0.70–0.76); validation 75.29% (CI 0.74–0.76) vs 65.93% (CI 0.64–0.67), even in multi-ethnicity populations. The clinical implications of the new risk score include prioritization of risk factor modification and personalized care at the patient level with improved targeting of interventions for stroke prevention.