Showing posts with label research. Show all posts
Showing posts with label research. Show all posts

Wednesday, August 24, 2011

Size Matters - 機器學習的分散式和平行處理議題

John Langford的個人部落格看到這個消息:由 Ron Bekkerman (LinkedIn),John Langford (Yahoo! Research)和 Misha Bilenko(Microsoft Research)共同編輯的 Scaling up Machine Learning 將在今年底出版。而且他們將在 KDD 2011 以這本書為基礎發表 Scaling Up Machine Learning 的 Tutorial

依照 John Langford 的介紹:
This tutorial focuses on providing an integrated overview of state-of-the-art platforms and algorithm choices. These span a range of hardware options (from FPGAs and GPUs to multi-core systems and commodity clusters), programming frameworks (including CUDA, MPI, MapReduce, and DryadLINQ), and learning settings (e.g., semi-supervised and online learning). The tutorial is example-driven, covering a number of popular algorithms (e.g., boosted trees, spectral clustering, belief propagation) and diverse applications (e.g., speech recognition and object recognition in vision).

不在現場的我,是沒有機會親聆這場盛會啦,不過稍微瞄了下簡報檔,覺得這本書裡面的題材都蠻有意思的,比如說下面第一個圖提到克服隱私權疑慮的嘗試,和第二個圖 Tree Ensembles 的說明。

期待這本書的上市,不過看到價格實在是讓人有點遲疑啊,哈哈!




Sunday, April 10, 2011

失之毫釐是不是謬以千里

在西班牙電信公司 Telefonica 研究院 工作的學者 Xavier Amatriain,前幾天在網誌上發表了一篇文章 Recommender Systems: We're doing it (all) wrong ,談到研究推薦系統的學者和開發者,在使用數據時,務必要注意數據的性質。

很多人使用 Likert Scale 做評分(Ratings)的量表基礎,比如說像「非常不喜歡、喜歡、無所謂、不喜歡、非常不喜歡」這樣的評分表就極爲常見,但是 Xavier 提醒我們 Likert Scale 的數據是 ordinal data ,這種數據僅僅表達次序關係,但是兩兩評分之間未必是 equidistant 的。若用這樣的數據計算距離(計算距離是相似性的基礎),其結果可能是失真的,循此邏輯推演下去,計算推薦系統準確率的指標 RMSE 的意義也可能失準。

從數學的角度來看,誤用定義當然是極爲嚴重的基本功的失誤,但是若從實務上考量,把 Likert 式評分當做 internal data,對推薦系統的成果究竟影響又多大,實在不好說 。不過,看來在這一點上不察,誤把馮京當馬涼的研究人員和開發人員可能不少哦!

Xavier Amatriain 寫這篇文章,是受 Judy Robertson 在 Blog@ACM 上的文章 We're Doing It Wrong 所啓發。Judy 在文中提到 2010 ACM Conference on Human Factors in Computing Systems  有學者發表研究 前一年會議中發表論文《Powerful and consistent analysis of Likert-type rating scales 》,爬梳學者使用的數據和統計工具,發現驚人的事實,原文是這樣的:
Kaptein, Nass, & Markopoulos (2010) published a paper in CHI last year found that in the previous year's CHI proceedings, 45% of the papers reported on likert type data but only 8% used non-parametric stats to do the analysis. 95% reported on small sample sizes (under 50 people). This is statistically problematic even if it gets past reviewers!

使用 Likert Scale 作爲實驗分析方法的學者竟然約略達到五成,Judy 在文章下半部提出她對此現象原因的觀察和建議,我對統計是大外行,只能點頭諾諾。但最抓住我眼球的句子是“95% reports on small sample size”這句,產業界鮮少有人信服學界真能做出「有用」的東西,確實有點道理,怨不得人。


[參考資料]
Kaptein, M., Nass, C., Markopoulos, P. (2010) Powerful and consistent analysis of Likert-type rating scales. In Proceedings CHI 2010, ACM, New York, NY, 2391-2394. DOI= http://doi.acm.org/10.1145/1753326.1753686

Saturday, February 5, 2011

Saturday, January 15, 2011

三杯通大道,一斗合自然



又一個舉杯的理由,薄酒不僅能讓墨客「酒發雄談,劍增奇氣,詩吐驚人語」,還能提升科研成果,突破瓶頸。「方知一杯酒,猶勝百家書」,古人誠不我欺!

日本國立材料科學院,一個研究超導體的研究團隊,在某次慶功小聚中,喝的有點多的研究人員,酒意上頭之後,決定把研究材料放到“許多、許多”酒 (原文是 many many liquor)中。

事後檢測,這批泡過慶功酒的材料,傳導性比平日的實驗結果好得多 (疑問:這些人原本要慶祝什麼?)。 進一步比較後發現,日本燒酒的成績比平日材料好 23%,泡過紅酒的材料則改善了 62%。但是根據報導,這宴會還供應威士忌和啤酒,莫非這兩種酒的表現不大好!?

我想,以嚴謹為尚的研究人員,應該還要比較混合不同酒類比例,以及不同廠牌的影響才是。

不論如何,一定很多人愛死這篇報導的結論了,”So, a little sip of something turns out to make potential superconductors much better at their jobs. And, perhaps, scientists better at their jobs as well.“。

三杯通大道,一斗合自然,如果喝一斗,就能生一篇SCI文章,....

Thursday, June 24, 2010

複習: Social Network Sites 的定義

近日在讀一篇有關 mobile social network 的文章,於是找出 danah m. boydNicole B. Ellison 寫的 Social Network Sites: Definition, History, and Scholarship,從頭複習 social network sites 的定義和觀念。這篇文章是為 Journal of Computer-Mediated Communication  2007年10月出版的社群網路特刊寫的,是當期特刊的導讀。

danah boyd 和 Nicole Ellison 這兩位對社群網路服務的所下的定義是以用戶的基本資料(profile)為核心,首先每個用戶必須在該網站建立一個用戶檔案,這個檔案的內容與資料開放程度隨網站的服務內容、方式和隱私政策而有所不同。

但是不管用戶檔案包含多少東西,每個用戶都必須建立與維護一個列表(list),這個列表包含所有在該網站生態系統裡與此用戶發生關聯(connection)的其他用戶。最重要的是,每個用戶都能瀏覽自己及其他用戶所建立的列表,社群網路的各式服務及其創意與變化的基礎,就是奠基在這個作者所強調的第三點。因為能看到與自己有關聯的用戶的社會關係列表,因此可以產生新的關聯,產生關聯(connection、relationship)之後,管理這些關聯的規則,和形成關聯後的用戶間可以執行的動作(actions),就是每個社群網路網站服務特色之所繫了。

文章中關於 social network sites  定義的原文如下,一個原本大家都以為知道「那是什麼」的一件事,真要寫出定義,還真是不簡單。難怪有人說 ...(以下消音).....

We define social network sites as web-based services that allow individuals to (1) construct a public or semi-public profile within a bounded system, (2) articulate a list of other users with whom they share a connection, and (3) view and traverse their list of connections and those made by others within the system. The nature and nomenclature of these connections may vary from site to site.

本文關於定義說明有一點頗值得琢磨,作者說雖然在網路上許多人把 social network sites 和 social networking sites 交替使用,視為同一個概念,但是作者認為 networking 更強調社群網路中與陌生人發展關係(initiate relationships)的現象,作者認為發展關係不是社群網路的全部,所以她們認為 social network sites 這個詞組更能表達社群網路概念所涵蓋的範圍。

為什麼 networking 不涵蓋管理 connected nodes 間關係的動作,只適宜描述管理 strangers 間新關係的建立管理等動作?這個說法不是很有說服力,至少沒有說服我。

Monday, April 26, 2010

[Video] Damon Horowitz at TEDxSoMA - Why Machines Need People

推友@alisohani 強力推薦 Aardark  共同創辦人 Damon Horowitz 今年一月在 TEDxSoMa 的一場演講 Why Machines Need People,TEDxSoMa 網站簡明扼要的說明 Damon 了他進入職場後的精彩轉折,同時也在文中說明了演講的重點。這是一個放下職場光環,回到學校念哲學博士,再回職場創業的牛人,這裡不是說有 PhD 學位的人都叫做哲學博士的那個“哲學”,他是真的到 Stanford 主修哲學!

這個當初對於技術懷抱宗教情懷 (when I went into college, I became very religious),篤信技術(I believe in technology)威力的牛人,在這個演講裡給我們的 final take away 是 Technology cannot solve all of our problems for us; the task of thinking is still ours。有意思,是不?

補充一點,Damon 說話的速度很快,還好Youtube 的字幕功能已經進化到夠強悍的地步。



Saturday, April 17, 2010

What makes a good paper?

IBM Almaden 研究中心 的研究員 Tessa LauACM 通訊部落格發表了一篇文章,談到幾個學術會議的論文審稿過程的不足與爭議,文章末了她提出她心目中合格的 HCI  論文,應該要達到的標準,雖然她談的是她專長的  HCI 領域, 我認為這些標準對於資工領域的研究都是適用的,所以抄錄於後:

  • Clear and convincing description of the problem being solved. Why isn't current technology sufficient? How many users are affected? How much does this problem affect their lives?
  • How the system works, in enough detail for an independent researcher to build a similar system. Due to the complexities of system building, it is often impossible to specify all the parameters and heuristics being used within a 10-page paper limit. But the paper ought to present enough detail to enable another researcher to build a comparable, if not identical, system.
  • Alternative approaches. Why did you choose this particular approach? What other approaches could you have taken instead? What is the design space in which your system represents one point?
  • Evidence that the system solves the problem as presented. This does not have to be a user study. Describe situations where the system would be useful and how the system as implemented performs in those scenarios. If users have used the system, what did they think? Were they successful?
  • Barriers to use. What would prevent users from adopting the system, and how have they been overcome?
  • Limitations of the system. Under what situations does it fail? How can users recover from these failures?


Monday, April 5, 2010

[簡報] 雲端應用於數位典藏的思考

ilyagram2010 網際網路趨勢研討會中演講,談他對雲端應用於數位典藏的思考,簡報內容可從研討會網站下載,也可至 SlideShare 瀏覽。簡報內容簡潔有力,條理清楚,是很好的示範,簡報第20頁,有簡報中使用圖檔 Credit 說明,是國內簡報比較少見到的,這一點很值得學習。

Saturday, March 27, 2010

Daniel Lemier's advice on How to Write Good Papers

Daniel Lemire 是我很敬佩的一位學者,他在推薦系統領域知名度相當高,而且是學術界投入部落格書寫的先行者之一。他的部落格並不局限於推薦系統議題,還涉及演算法和資訊科學理論的基礎問題,他還不時以自身的經驗提供出如何做研究、寫論文的建議,文章品質都很棒。昨日 Daniel Lemire 在 SlideShare 放上他針對研究生演講(talk?)如何撰寫論文的簡報,看了簡報,我只有一個想法,我實在是太混了。


Friday, March 26, 2010

[PDF] How to read a research paper

哈佛大學教授 Michael Mitzenmacher (部落格 My Biased Coin 的作者)提供一份很棒的如何閱讀論文 (下載電子檔) 的指引和建議,並且願意提供原稿(他是用 Latex 寫的),讓大家繼續發揚改進。Mitzenmacher 教授在他的部落格還特別推薦Jason EisnerHow to Read a Technical Paper 的最後一節 What to read



What to read

  • creative web search



    • experiment with several searches
    • put yourself in an author's shoes; what phrases might they have used?
    • become a power searcher! (read the help pages for your search engine) 

  • find related work



    • backward references: follow the bibliography to earlier papers
    • forward references: see who else has cited the work (via an interface such as Google Scholar




    • has someone else already listed the right papers for you?



      • survey papers in journals (also called "review articles")
      • course syllabi
      • reading group webpages
      • chapters in textbooks
      • online tutorials
      • literature review chapters from dissertations
      • direct recommendations from friends or professors (perhaps at other institutions) 




      • breadth-first exploration



        • read a lot of abstracts (and skim the papers as needed) before deciding which papers are best to read
        • it's okay to read multiple related papers at once, flipping back and forth so that they clarify one another
        • to get a feel for the research landscape in an area, flip through the proceedings of a relevant recent workshop, conference, or special-theme journal issue 




        • when the going gets tough, switch to background reading



          • textbooks or tutorials
          • review articles
          • introductions and lit review chapters from dissertations
          • early papers that are heavily cited
          • sometimes Wikipedia

        Tuesday, March 2, 2010

        Is "suffering" an indispensable part of research?

        看到 Research as a second language 談 "mentoring .vs. coaching" 的文章,忍不住要問 Is "suffering" an indispensable part of research? 

        Mentoring vs. Coaching
        The Centre for Development of Human Resources and Quality Management in Denmark is holding a conference to present the results of a PhD coaching project. The project, which involved PhD students from three universities, appears to have been a success.

        Specifically, they discovered that:
        • The participants got a lot out of the coaching.
        • The coaching did not get in the way of traditional academic supervision
        • Individual coaching works better than workshops.
        My experience confirms these conclusions. But the theme of the conference appears to be captured in the question, "Do you necessarily have to go through a lot of suffering to get a PhD?" That question, and the fact that the coaching was found "not to disturb" academic supervision, got me thinking about what I do. In fact, it got me reconsidering.

        First, I do believe that "suffering" is an important part of research (in Danish, as Kierkegaard pointed out, suffering rhymes with science). Second, I've long noticed that "coaching" is often an inappropriate metaphor because outside of actual sports the "coach" is often not a master of the craft she coaches; rather she has a generalized ability to motivate others and help them get organized. (This, by the way, does not always mean she has an ability to get herself organized or get anything done herself.) Like me, she may not know very much about the area of scholarship that the PhD student is working in.

        .......


        Tuesday, February 16, 2010

        NP 問題就這樣解決了!?

        P .vs. NP 是計算機科學一個很重要的未解問題,等級為 P 的問題是能在多項式時間內解決的問題,NP 問題則是能在多項式時間驗證解答是否正確的問題,顯然 P 問題是 NP 問題的子集,但 P 是否等於 NP,一直無法證明,也無法推翻

        財務領域也有個很重要的有效市場假說efficient market hypothesis),在有效市場裡,市場價格已經充分反應所有證卷市場的歷史訊息,所以分析過去的資料以預測市場未來走向是行不通的。

        這兩個問題似乎是八杆子打不着關係,但是有位牛人Phil Maymin)嘗試證明這兩個問題是等價的,他投到 arxiv 的論文說他能夠證明如果市場是弱式有效市場(weak-form efficient market),那麼 P = NP,反之亦然。到目前為止,這篇文章我只讀了三頁,我想就算通篇讀完,以我的程度也無法判斷這位牛人說的究竟有沒有道理。若有高人能現身釋疑,小弟感激不盡。總之,我的感想,這位多才多藝的牛人,實在是太牛了

        這篇文章應該要歸檔至 to_read_later 類別,只是 later 究竟是多久以後,我也不知道?

        Markets are efficient if and only if P = NP
        Abstract: I prove that if markets are weak-form efficient, meaning current prices fully reflect all information available in past prices, then P = NP, meaning every computational problem whose solution can be verified in polynomial time can also be solved in polynomial time. I also prove the converse by showing how we can "program" the market to solve NP-complete problems. Since P probably does not equal NP, markets are probably not efficient. Specifically, markets become increasingly inefficient as the time series lengthens or becomes more frequent. An illustration by way of partitioning the excess returns to momentum strategies based on data availability confirms this prediction.


        Thursday, February 4, 2010

        大丈夫當如是另解

        John Battelle (The Search 的作者) 昨日撰文推介 Aardvark 公司的  Damon HorowitzSepandar D. Kamvar 一篇將在 WWW 2010 發表的論文 - Anatomy of a Large-Scale Social Search Engine,這篇文章和 Sergey Brin 和 Larry Page 撰寫的 PageRank 基礎的那篇著名論文只差一個字,這當然不是巧合,因為這篇文章不僅介紹 Aardvark 所開發的搜尋引擎, 並比較他們的做法和傳統搜索引擎的不同。就如 John Battelle 所介紹的,這兩位創業者的企圖心相當驚人,不僅試圖把自己的作品和成就 Google 的里程碑”接軌“,而且在矽谷,在創業前期的公司敢大方公開自己的技術基礎的例子實在不多。

        John Battelle 條理分明的介紹這篇論文有趣和有意義(worth a read)的地方,論點十分有說服力,有興趣的人可以直接找原文來看,不過 John 也警告大家這篇論文有很多數學,大家自己看着辦(grin)。但是敝人覺得 John 的文章中最大的亮點是他提到 Brin & Page 的文章在 Google 創業之前並沒有一炮而紅,而是...couldn't even get their paper accepted。想到像 Brin & Page 這樣的大才尚且有這般經歷,似我等這般凡夫俗子,在投稿歷程上屢敗屢戰也不是什麼稀罕的事了。想到這裡,連日陰霾的心情似乎有點好轉...
        First, the paper's authors, both of whom have worked at Google, clearly have a sense of potential history here, in that they not only crib Google's original paper's title, they also mirror the first line (substituting "Aardvark" for "Google", of course). Now that's some b*lls. Of course, when Larry and Sergey first presented Google, they couldn't even get their paper accepted (it took three tries, if I recall correctly. Someone should write a book about that...).

        John Battelle 認為這件事是個寫書的好點子,what do YOU think?

        Monday, June 29, 2009

        細節是魔鬼還是機會

        搜尋產業專家Danny Sullivan 在Michael Jackson 驟逝之後,細心觀察網際網路的動態以及重要的搜尋引擎在面對這個事件的表現,發現 Google 曾經犯了一個很離譜的錯誤,搜尋 michael jackson dead ,竟然會得到 MJ 死於 2007 享年65 歲的結果 (這個發現又得歸功於於無處不在的推友)。


        這個錯誤當然很快就被發現而且更正了,Danny Sullivan 解釋發生這錯誤顯然是因為谷歌在使用 Wikipedia 資料時,錯用了另外一位 Michael Jackson 的資料。Matthew Hurst 認為這是一個典型的在執行細節上失誤的範例,並且仔細的分析執行文字探勘(text mining)工作時每一個步驟的細節,還語重心長的說 Attention to detail will always be a killer feature!

        最後,筆者忍不住要加上一段有點 cynical 的按語:身為微軟員工的 Matthew Hurst 在分析谷歌這次失誤的時候,是怎樣的心情咧!?

        Thursday, April 30, 2009

        [Updated] Blogs on Data Mining, Web Mining and Analytics

        [ Last Updated : 2009/04/30]
        The original list was compiled by KDnuggets, I added some my collections and put the modified list onto the blog of my own. The list is growing and changing all the time, I will keep on maintaining my own list continually(2007/04/25).

        Avinash Kaushik published his Top Ten Web Analytics Blogs: July 2007 on July 27, I added his Top Ten to my list as well (2007/07/29).

        I added some more blogs to THE list(2007/08/15).

        Added Coremark Analytics and Life Analytics (2007/08/23).

        Added DM(X) (2007/11/26).

        Added Beyond Search and Daniel Lemire's Blog (2007/12/31).

        Added IR Thoughts, Herself's Artificial Intelligence, Just a guy in garage and Datawocky (2008/04/01)

        Data Mining Research maintained a big list of data mining blogs. I combined his collection and my original list (2008/05/21).

        Added Radford Neal's Blog and A blog by Tim Manns.(2008/09/01).

        Added Datalligence (2008/10/09).

        Added Monash University Business Intelligence Blog , The Noise Channel (2008/11/03)

        Added Dynamic Notions (2008/11/04)

        Added Jeff's Search Engine Caffe, Statistical Modeling, Causal Inference, and Social Science, Brendan O’Connor’s Blog - AI and Social Science (2008/12/4)

        Smart(Enough) Systems renamed to James Taylor on Enterprise Decision Management and the link pointed to a new address (2008/12/16)

        Added anuradha@NumbersSpeak (2009/05/01)

        First posted on Apr 25,2007

        Friday, January 23, 2009

        Question: Netflix Prize within months?

        獲得2008年度 Netflix Progress Prize 的團隊 BellKor in BigChaos 在今年一月初,又提交一次成果,這回他們把成績從 RMSE 0.8616 推進到 RMSE 0.8598,距離得獎所需的標準 0.8563 又推進一大步。


        雅虎的資深研究員 John Langford 在閱讀 BellKor in Chaos 釋出的文件(1,2)後,在他的個人部落格 Machine Learning (Theory) 指出也許距離獎金揭曉的時間不遠了,他同時指出該團隊演算法包含了 stochastic gradient descent, ensemble prediction, and targeting residuals 各領域的技術,並且指出他們在 2008 年間將演算法參數化的努力。他同時意味深長的說,或許- the right parameterization might very well succeed - 正確的參數就能將大獎抱走呢!
        Several aspects of solutions are taken for granted including stochastic gradient descent, ensemble prediction, and targeting residuals (a form of boosting). Relatively to last year, it appears that many approaches have added parameterizations, especially for the purpose of modeling through time.
        當然,John Langford 也在他的短文裡,提出了他的疑慮:One fear is that the progress is asymptoting on the wrong side of the 10% threshold,可是他也承認從去年底到今年一月的進步的的確令人印象深刻 。

        不過我個人比較好奇的是這些演算法,能不能作適當的改變後,應用到其他的產品,或許參數化正是往這個方向努力的指標之一,但是我對 overfitting 這件事仍然有點疑慮。雖然還沒有時間深入研讀相關文件,但在 John Langford 的部落格留言區以及 Netflix 網站的論壇裡有讀者討論到 overfitting 的問題,雖然有人認為這個比賽已經在各方面取得平衡(請參考這裡),不需擔心 overfitting,但是我個人仍然存疑。

        或許 - Time will tell。

        Tuesday, December 16, 2008

        Reading List: Diversity in Recommenders

        Daniel Lemire 在上個月整理他認為與推薦系統的多元推薦輸出(diversity of recommendation lists)有關的文獻,有些讀者在留言裡也提出他們的建議。初步過濾之後,我把自己感興趣的文章,用 CiteULikeRefworks 的輸出功能,製作IEEE 格式書目如後,作為備忘查考之用:

        [1] C. Clarke, M. Kolla, G. Cormack, O. Vechtomova, A. Ashkan, S. Büttcher and I. Mackinnon, "Novelty and diversity in information retrieval evaluation," in SIGIR '08: Proceedings of the 31st Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, 2008, pp. 659-666.

        [2] D. Fleder and K. Hosanagar, "Blockbuster Culture's Next Rise or Fall: The Impact of Recommender Systems on Sales Diversity," SSRN eLibrary, 2008.

        [3] D. Fleder and K. Hosanagar, "Recommender systems and their impact on sales diversity," in EC '07: Proceedings of the 8th ACM Conference on Electronic Commerce, 2007, pp. 192-199.

        [4] L. Iaquinta, M. de Gemmis, P. Lops, G. Semeraro, M. Filannino and P. Molino, "Introducing Serendipity in a Content-Based Recommender System," Hybrid Intelligent Systems, 2008. HIS '08. Eighth International Conference on, pp. 168-173, 2008.

        [5] Q. Le and A. Smola, "Direct Optimization of Ranking Measures," Apr 2007. [Online]. Available: http://arxiv.org/abs/0704.3359.

        [6] D. Lemire, S. Downes and S. Paquet, "Diversity in open social networks," 2008.

        [7] L. Mcginty and B. Smyth, "On the Role of Diversity in Conversational Recommender Systems," 2003.

        [8] S. Mcnee, J. Riedl and J. Konstan, "Being accurate is not enough: How accuracy metrics have hurt recommender systems," in CHI '06: CHI '06 Extended Abstracts on Human Factors in Computing Systems, 2006, pp. 1097-1101.

        [9] K. Swearingen and R. Sinha, "Beyond algorithms: An HCI perspective on recommender systems," 2001.

        [10] Y. Xu and H. Yin, "Novelty and topicality in interactive information retrieval," J. Am. Soc. Inf. Sci. Technol., vol. 59, pp. 201-215, 2008.

        [11] C. Zhai, W. Cohen and J. Lafferty, "Beyond independent relevance: Methods and evaluation metrics for subtopic retrieval," in SIGIR '03: Proceedings of the 26th Annual International ACM SIGIR Conference on Research and Development in Informaion Retrieval, 2003, pp. 10-17.

        [12] F. Zhang, "Research on Recommendation List Diversity of Recommender Systems," Management of e-Commerce and e-Government, 2008. ICMECG '08. International Conference on, pp. 72-76, 2008.

        [13] M. Zhang and N. Hurley, "Avoiding monotony: Improving the diversity of recommendation lists," in RecSys '08: Proceedings of the 2008 ACM Conference on Recommender Systems, 2008, pp. 123-130.


        如果需要下載這些文章的電子檔,請到筆者的 CiteULike 資料庫(tag: Diversity)查看論文下載位址的細節資料。

        Monday, June 30, 2008

        When all models are unnecessary, ...

        "All models are wrong, and increasing you can succeed without them."

        最近一期 Wired 的封面故事:The End of Theory: The Data Deluge Makes the Scientific Method Obsolete,標題驚人,內容頗具爭議性,在部落圈和學術界掀起一陣討論和討伐之風。

        這篇文章談到資料挖掘在Google 的成功中扮演的角色,以及可能在未來科學研究中扮演的角色,企圖雄偉,但是立論薄弱,而且對有些基本的東西有誤解,所以文章一經發表,讀者反應激烈,用句俗諺來形容,可以說是捅了馬蜂窩

        Wired 網站裡的讀者回應區,立刻有人反應作者做了能力範圍(out of league)外的事情,也有人認為他越線(crossed the line)了,甚至有位密西根大學的教授(Cosma Shalizi)在自己的部落格說出 I recently made the mistake of trying to kill some waiting-room time with Wired 的狠話。

        這篇文章由Wired 雜誌主編 Chris Anderson長尾理論的發明者)執筆,Chris Anderson 的確不愧為暢銷書作者,文采斐然沒有話說,先以統計學家 George Box 的著名警句 All models are wrong, but some are useful. 破題,然後以優美的排比句子,揭示 Petabyte 時代的來臨:

        Sixty years ago, digital computers made information readable. Twenty years ago, the Internet made it reachable. Ten years ago, the first search engine crawlers made it a single database.

        Petabyte Age 不是文人的夸飾,隨著資訊科技的進步,人類累積和儲存資料的本事越來越大,今年初(January 2008),Google 發表的 MapReduce 論文,透露了 Google 一天要處理 20 Petabytes 的資料。大量的數據,加上資料挖掘以及統計的幫助,讓 Google 的競爭力如虎添翼,谷歌本身的成就和他們贊助的生物資訊研究,充分說明了資料(數據)的重要性。所以 Peter NorvigGeorge Box 的名句改成 All models are wrong, and increasingly you can succeed without them.

        之前筆者也曾撰文討論過資料在數據挖掘研究裡的重要性,但是 Chris Anderson 在這裡走進了推演的誤區,把資料的重要性無限上綱,得到了只要有大量數據和應用數學(applied math;顯然他想說的是 data mining),天下沒有辦不到的事。甚至他認為這是 paradigm (有人翻譯為範式,還有更好的翻譯嗎?)的轉移,所以才有 End of Science 這樣驚人的標題。簡而言之,他認為老套的做學問的方法過時了:

        It's science. The scientific method is built around testable hypotheses. These models, for the most part, are systems visualized in the minds of scientists. The models are then tested, and experiments confirm or falsify theoretical models of how the world works. This is the way science has worked for hundreds of years.

        ..... omitted....

        Once you have a model, you can connect the data sets with confidence. Data without a model is just noise.  But faced with massive data, this approach to science — hypothesize, model, test — is becoming obsolete.

        他進一步闡述他的看法,有了數據、電腦、演算法,把數據丟進運算機器之後,我們只要等待結果就行了,不需要假設、模型,也不需要相關的知識。就像 Google 應用統計結果做機器翻譯和拼字檢查,不需要懂語言,也能得到很棒的結果:

        There is now a better way. Petabytes allow us to say: "Correlation is enough." We can stop looking for models. We can analyze the data without hypotheses about what it might show. We can throw the numbers into the biggest computing clusters the world has ever seen and let statistical algorithms find patterns where science cannot.

        Chris Anderson 所描述的新科學,在論述上並不充分完整。誠然,資料+運算能力+資料挖掘演算法的公式,在可以產生大量資料的領域,例如:太空、物理、製藥、基因等等,可以得到很棒的結果。但是,並不是所有的研究領域都有這樣的條件(可收集大量數據),這個新公式是否無往不利,包山包海,是很大的疑問。再者,不同領域的研究方法論也不同,貿然進入結論,認為新公式就是新的科學典範,是太草率了。比如說,有人質疑,如果這個模式成立,我們如何發現新的東西,因為我們不知道「新」東西的資料要從哪裡來?

        這種把科學的過程目的過度簡化的 "science without model & correlation supersedes causation" 理論,是有大問題的,Chris Anderson 所描述的新科學 — 從大量資料裡發現值得注意的資訊或知識 ,只是知其然的地步,這才是科學的起步而已;科學家的使命是知其所以,要知道事物的細節和所有發生現象的解釋,才是推動科學(和科技)進步的動力。

        John Timmer (ars technica) 在他的文章裡面,問了一個有趣的問題,如果一個理論不能提供可驗證的假設(testable hypotheses),我們怎麼知道我們錯的有多嚴重?推翻 testable hypotheses 的的必要性,我們甚至不知道結果是對是錯?

        而且 Chris Anderson 的說法也很容易誤導大家理解資料挖掘的真意,每個執行過資料挖掘工作的人(不論是學者、分析師、工程師)都知道,在找出規則(rules)和型樣(patterns)之後,如何判斷找出的資訊是否有用、有效,正需要上面所述知其所以的能力,才能充分利用挖掘出的資訊得到最大效益。資料挖掘絕對不只是 number crunching 的黑箱,沒有理論,沒有假設,沒有關聯,是不可能完成一個資料挖掘任務的。

        所以我完全同意 John Timmer 所說的: At a more fundamental level, in spite of what Chris Anderson has to say, science is about explanations, coherent models and understanding。他對關聯(correlations)和模型的解釋,更是簡明有力,深得我心:

        Correlations are a way of catching a scientist's attention, but the models and mechanisms that explain them are how we make the predictions that not only advance science, but generate practical applications.

        在賓州大學任教的 Fernando Pereira 針對這篇文章的評論也很具參考價值:

        I like big data as much as the next guy, but this is deeply confused. Where does Anderson think those statistical algorithms come from? Without constraints in the underlying statistical models, those "patterns" would be mere coincidences. Those computational biology methods Anderson gushes over all depend on statistical models of the genome and of evolutionary relationships.

        Those large-scale statistical models are different from more familiar deterministic causal models (or from parametric statistical models) because they do not specify the exact form of observable relationships as functions of a small number of parameters, but instead they set constraints on the set of hypotheses that might account for the observed data. But without well-chosen constraints — from scientific theories — all that number crunching will just memorize the experimental data.

        個人認為當Chris Anderson 說出 Forget taxonomy, ontology and psychology 時,顯然是走得太遠,有點忘形了。雖然從作文章的角度來看,這在修辭上是很有講究的句子,但是這些文字透露了推理的輕率和治學態度的傲慢,我想這是眾多部落格作者和學者看不過去的原因之一。

        更多數據對於科學家絕對是好事,越多數據越能驗證假設的正確性,但是更豐富充足的數據絕對不代表我們就可以揚棄「大膽假設小心求證」的治學原則了。不過,整體而言,Chris Anderson 的說法倒也不是全無道理,Google 的成功方程式對於學術界還是有一定影響的,Kevin Kelly 在文章裡引用 George Dyson說法很值得參考:

        What Chris Anderson is hinting at is that Science (and some very successful business) will increasingly be done by people who are not only reading nature directly...,They accomplish what science does, although not in the traditional manner...

        更多的數據讓科學家們多了一種有別于傳統的工作方式,但並不代表傳統的終結, John Timmer 的結語說得好:

        Overall, the foundation of the argument for a replacement for science is correct: the data cloud is changing science, and leaving us in many cases with a Google-level understanding of the connections between things. Where Anderson stumbles is in his conclusions about what this means for science. The fact is that we couldn't have even reached this Google-level understanding without the models and mechanisms that he suggests are doomed to irrelevance.

        除了以上整理的觀點之外,也有像 Matthew Hurst 這樣保持冷靜的作者,雖然他也對 Chris Anderson 的文章不滿,但他希望在整理出完整而有意義的想法之後,再發表自己的意見,讓我們拭目以待吧...

         

        (Strongly recommended : 我在 Diigo上建了一個 List: End of Theory,相關參考資料都加入我的資料庫裡了,讀者可以閱讀這個 WebSlide,瀏覽資料庫的內容。)

        參考資料:

         

        Share this post :

        如果我的心是一朵蓮花

        ~ 林徽因 · 馬雁散文集 · 蓮燈 ~ 馬雁 在她的散文《高貴一種,有詩為證》裡,提到「十多年前,還不知道林女士的八卦及成就前,在期刊上讀到別人引用的《蓮燈》」 覺得非常喜歡,比之卞之琳、徐志摩,別說是毫不遜色,簡直是勝出一籌。前面的韻腳和平仄的處理顯然高於戴...