顯示具有 資料探勘 標籤的文章。 顯示所有文章
顯示具有 資料探勘 標籤的文章。 顯示所有文章

7/13/2025

Statistical Modeling: The Two Cultures

Cynthia Rudin, Leo Breiman, the Rashomon Effect, and the Occam Dilemma, arXiv:2507.03884, 2025.

In the famous “Two Cultures” paper, Leo Breiman provided a visionary perspective on the cultures of “data models” (modeling with consideration of data generation) versus “algorithmic models” (vanilla machine learning models). I provide a modern perspective on these two approaches. One of Breiman’s key arguments against data models is what he called the “Rashomon Effect,” which is the existence of many different-but-equally-good models. The Rashomon Effect implies that data modelers would not be able to determine which model generated the data. Conversely, one of his core advantages in favor of data models is simplicity, as he claimed there exists an “Occam Dilemma,” i.e., an accuracy-simplicity tradeoff, where algorithmic models must be complex in order to be accurate. After 25 years of more powerful computers, it has become clear that this claim is not generally true, in that algorithmic models do not need to be complex to be accurate; however, there are nuances that help explain Breiman’s logic, specifically, that by “simple,” he appears to consider only linear models or unoptimized decision trees. Interestingly, the Rashomon Effect is a key tool in proving the nullification of the Occam Dilemma. To his credit though, Breiman did not have the benefit of modern computers, with which my observations are much easier to make.

2/02/2023

學習大數據 (big data) 的技能

一些工具或念個學位

可以參考 DS Examiner, Data Scientist Foundations: The Hard and Human Skills You Need, November 8, 2013

或者  Insight Data Science Fellows Program 說明了可能使用的工具
  1. Software Engineering Best Practices: Learn how to contribute to a large code-base and instrument a web application to collect data. Tools you may learn: Python, Git, LAMP web stack, Javascript, Flask.
  2. Storing and Retrieving Data: How to clean data, store it in the appropriate database or distributed data storage system and then run queries to retrieve the information needed for analysis. Tools you may will learn: MySQL, Hadoop, Hive.
  3. Statistical Analysis & Machine Learning: Learn industry best practices for doing basic and advanced statistical analysis on large data sets. Tools you may learn: R, NumPy & SciPy, Mahout.
  4. Visualizing and Communicating Results: Learn how to effectively communicate your findings visually and verbally. Tools you may learn: D3 Javascript library, visualization and presentation best practices. 

9/15/2022

Cynthia Rudin wins AAAI Squirrel AI Award

AAAI, Duke Computer Scientist Wins $1 Million Artificial Intelligence Prize, A 'New Nobel', October 12, 2021.

Whether preventing explosions on electrical grids, spotting patterns among past crimes, or optimizing resources in the care of critically ill patients, Duke University computer scientist Cynthia Rudin wants artificial intelligence (AI) to show its work. Especially when it’s making decisions that deeply affect people’s lives.

2/21/2022

7 real-world applications of reinforcement learning

 Joy Zhang, 7 real-world applications of reinforcement learning, gocoder, February 17, 2022

1. Autonomous driving with Wayve

2. Personalizing your Netflix recommendations

3. Optimizing inventory levels for Walmart

4. Improving search engine results with search.io

5. Improving language models with OpenAI's WebGPT

6. Trading on the financial markets with IBM's DSX platform

7. Robotics with the University of California, Berkeley

10/18/2021

Minimum-Distortion Embedding

Akshay Agrawal, Alnur Ali and Stephen Boyd (2021), "Minimum-Distortion Embedding", Foundations and Trends® in Machine Learning: Vol. 14: No. 3, pp 211-378. http://dx.doi.org/10.1561/2200000090. 

We consider the vector embedding problem. We are given a finite set of items, with the goal of assigning a representative vector to each one, possibly under some constraints (such as the collection of vectors being standardized, i.e., have zero mean and unit covariance). We are given data indicating that some pairs of items are similar, and optionally, some other pairs are dissimilar. For pairs of similar items, we want the corresponding vectors to be near each other, and for dissimilar pairs, we want the corresponding vectors to not be near each other, measured in Euclidean distance. We formalize this by introducing distortion functions, defined for some pairs of the items. Our goal is to choose an embedding that minimizes the total distortion, subject to the constraints. We call this the minimum-distortion embedding (MDE) problem.

This monograph is accompanied by an open-source Python package, PyMDE, for approximately solving MDE problems. Users can select from a library of distortion functions and constraints or specify custom ones, making it easy to rapidly experiment with different embeddings. Because our algorithm is scalable, and because PyMDE can exploit GPUs, our software scales to data sets with millions of items and tens of millions of distortion functions. Additionally, PyMDE is competitive in runtime with specialized implementations of specific embedding methods. To demonstrate our method, we compute embeddings for several real-world data sets, including images, an academic co-author network, US county demographic data, and single-cell mRNA transcriptomes.

6/26/2021

Interpretable predictive maintenance for hard drives

Maxime Amram, Jack Dunn, Jeremy J.Toledano, and Ying Daisy Zhuo, Interpretable predictive maintenance for hard drives, Machine Learning with Applications, Volume 5, 15 September 2021, 100042.

Existing machine learning approaches for data-driven predictive maintenance are usually black boxes that claim high predictive power yet cannot be understood by humans. This limits the ability of humans to use these models to derive insights and understanding of the underlying failure mechanisms, and also limits the degree of confidence that can be placed in such a system to perform well on future data. We consider the task of predicting hard drive failure in a data center using recent algorithms for interpretable machine learning. We demonstrate that these methods provide meaningful insights about short- and long-term drive health, while also maintaining high predictive performance. We also show that these analyses still deliver useful insights even when limited historical data is available, enabling their use in situations where data collection has only recently begun.

5/03/2021

Machine Learning Under a Modern Optimization Lens

Dimitris Bertsimas and Jack Dunn, Machine Learning Under a Modern Optimization Lens, Dynamic Ideas LLC, 2019.
The book provides an original treatment of machine learning (ML) using convex, robust and mixed integer optimization that leads to solutions to central ML problems at large scale that can be found in seconds/minutes, can be certified to be optimal in minutes/hours, and outperform classical heuristic approaches in out-of-sample experiments.

4/03/2021

How we use AutoML, Multi-task learning and Multi-tower models for Pinterest Ads

Ernest Wang, How we use AutoML, Multi-task learning and Multi-tower models for Pinterest Ads, Aug 20, 2020.

People come to Pinterest in an exploration mindset, often engaging with ads the same way they do with organic Pins. Within ads our mission is to help Pinners go from inspiration to action by introducing them to the compelling products and services that advertisers have to offer. A core component of the ads marketplace is predicting engagement of Pinners based on the ads we show them. In addition to click prediction, we look at how likely a user is to save or hide an ad. We make these predictions for different types of ad formats (image, video, carousel) and in context of the user (e.g., browsing the home feed, performing a search, or looking at a specific Pin.)

2/04/2021

台積電的數位轉型

王宏仁,台積電數位轉型的下一步,靠AI推動全面轉型(上),iThome,2021-01-27

IT,就是讓台積順利將各種製造服務轉為自動化的關鍵。在1996年,台積為了將後勤和財務資訊轉換成管理資訊來輔助決策,展開了資訊系統大升級,當年WWW技術才剛點燃了全球網際網路新浪潮不久,台積就能運用當時的網路技術,推出了全方位訂單管理系統,讓顧客透過電腦網路連線取得訂單和產品生產資訊,這也是半導體代工產業的創舉。1999年更推出供應鏈管理的資訊服務,可以讓客戶透過網路下單、即時查詢晶片生產進度和出貨狀況,早在30年前,台積電就採取了現在網路電商慣用的銷售形式。...

1/25/2021

Special Issue — M&SOM 20th Anniversary

Special Issue — M&SOM 20th Anniversary, Volume 22, Issue 1, January-February 2020 (online)

This special issue contains invited and review articles by eminent researchers in the field.

10/22/2020

50 years of Data Science

David Donoho, 50 Years of Data Science, Journal of Computational and Graphical Statistics, Volume 26, Issue 4, 2017, Pages 745-766.

More than 50 years ago, John Tukey called for a reformation of academic statistics. In “The Future of Data Analysis,” he pointed to the existence of an as-yet unrecognized science, whose subject of interest was learning from data, or “data analysis.” Ten to 20 years ago, John Chambers, Jeff Wu, Bill Cleveland, and Leo Breiman independently once again urged academic statistics to expand its boundaries beyond the classical domain of theoretical statistics; Chambers called for more emphasis on data preparation and presentation rather than statistical modeling; and Breiman called for emphasis on prediction rather than inference. Cleveland and Wu even suggested the catchy name “data science” for this envisioned field. A recent and growing phenomenon has been the emergence of “data science” programs at major universities, including UC Berkeley, NYU, MIT, and most prominently, the University of Michigan, which in September 2015 announced a $100M “Data Science Initiative” that aims to hire 35 new faculty. Teaching in these new programs has significant overlap in curricular subject matter with traditional statistics courses; yet many academic statisticians perceive the new programs as “cultural appropriation.” This article reviews some ingredients of the current “data science moment,” including recent commentary about data science in the popular media, and about how/whether data science is really different from statistics. The now-contemplated field of data science amounts to a superset of the fields of statistics and machine learning, which adds some technology for “scaling up” to “big data.” This chosen superset is motivated by commercial rather than intellectual developments. Choosing in this way is likely to miss out on the really important intellectual event of the next 50 years. Because all of science itself will soon become data that can be mined, the imminent revolution in data science is not about mere “scaling up,” but instead the emergence of scientific studies of data analysis science-wide. In the future, we will be able to predict how a proposal to change data analysis workflows would impact the validity of data analysis across all of science, even predicting the impacts field-by-field. Drawing on work by Tukey, Cleveland, Chambers, and Breiman, I present a vision of data science based on the activities of people who are “learning from data,” and I describe an academic field dedicated to improving that activity in an evidence-based manner. This new field is a better academic enlargement of statistics and machine learning than today’s data science initiatives, while being able to accommodate the same short-term goals.

10/01/2020

A few useful things to know about machine learning

Pedro Domingos, A Few Useful Things to Know About Machine Learning, Communications of the ACM, 2012, Vol. 55, No. 10, Pages 78-87.

This article summarizes 12 key lessons that machine learning researchers and practitioners have learned. These include pitfalls to avoid, important issues to focus on, and answers to common questions. 

Table 1 The three components of learning algorithms.

Learning = Representation + Evaluation + Optimization 

This is a nice overview article which could be assigned as an entry-level course reading. For a teacher, you could cover these topics or provide enough background material in your course so that the students could explore the content by themselves. 

9/18/2020

Introducing data science

Davy Cielen, Arno D. B. Meysman, and Mohamed Ali, Introducing data science: Big data, machine learning, and more, using Python tools, Manning, May 2016.

Introducing Data Science explains vital data science concepts and teaches you how to accomplish the fundamental tasks that occupy data scientists. You'll explore data visualization, graph databases, the use of NoSQL, and the data science process. You'll use the Python language and common Python libraries as you experience firsthand the challenges of dealing with data at scale. Discover how Python allows you to gain insights from data sets so big that they need to be stored on multiple machines, or from data moving so quickly that no single machine can handle it. This book gives you hands-on experience with the most popular Python data science libraries, Scikit-learn and StatsModels. After reading this book, you'll have the solid foundation you need to start a career in data science.

3/10/2020

佳世達集團讓 AI 產品化,解決 3 個餐飲業痛點

2020.03.09 
長年協助餐飲業客戶資訊數位化,益欣總經理康惠媚觀察,目前連鎖餐飲業客戶最需要的是會員管理、數位行銷、人力排班應用;而顧客的消費管道愈來愈多,留下很多足跡,業者也必須思考「人、貨、場」三方搜集的數據怎麼應用。...

11/08/2019

Nine Algorithms That Changed the Future (改變世界的九大演算法)


陳正芬譯改變世界的九大演算法:讓今日電腦無所不能的最強概念經濟新潮社2014
本書所介紹的九大演算法是:搜尋引擎的索引(search engine indexing)、網頁排序(page rank)、公鑰加密(public-key cryptography)、錯誤更正碼(error-correcting codes)、模式辨識(pattern recognition,如手寫辨識、聲音辨識、人臉辨識等等)、資料壓縮(data compression)、資料庫(databases)、數位簽章(digital signature),以及一種如果存在的話將會很了不起的偉大演算法,並探討電腦能力的極限。 
作者將我們日常生活會用到的電腦功能 背後的道理,以淺顯易懂的方式介紹,不具備資訊科學的背景也可以了解。而且令人驚喜的是,每一種演算法,都是一個解決問題的創意與線索,也讓我們得以一窺 近代數學家、資訊科學家的努力探索成果。面對越來越科技化的現代生活與職場挑戰,這些基本原理和概念值得我們去了解、吸收,為未來世界做好準備。

10/14/2019

電腦的計算速度和線性代數

為了說明電腦的計算速度,特別設計了一個內積的問題,不到一秒可以計算一千萬組數字內積。推薦系統有多種方法,方法之一使用內積 (inner product);另外,現代電腦可以快速地計算線性代數的問題,無形中培養計算思維 (Computational thinking),也可以廣泛地應用在國高中的教學。

import time
import numpy as np
t = time.time() # 現在系統時間
a = np.random.rand(10**7) # 產生 10**7 亂數
b = np.random.rand(10**7)
np.dot(a,b) # 內積
print("Jobs done in:", time.time()-t, " seconds") # 現在系統時間 減去 初始系統時間

配合 Google Colab,解決軟硬體不足的問題。