Saturday, January 21, 2017
Wednesday, August 24, 2011
Size Matters - 機器學習的分散式和平行處理議題
從John Langford的個人部落格看到這個消息:由 Ron Bekkerman (LinkedIn),John Langford (Yahoo! Research)和 Misha Bilenko(Microsoft Research)共同編輯的 Scaling up Machine Learning 將在今年底出版。而且他們將在 KDD 2011 以這本書為基礎發表 Scaling Up Machine Learning 的 Tutorial。依照 John Langford 的介紹:
This tutorial focuses on providing an integrated overview of state-of-the-art platforms and algorithm choices. These span a range of hardware options (from FPGAs and GPUs to multi-core systems and commodity clusters), programming frameworks (including CUDA, MPI, MapReduce, and DryadLINQ), and learning settings (e.g., semi-supervised and online learning). The tutorial is example-driven, covering a number of popular algorithms (e.g., boosted trees, spectral clustering, belief propagation) and diverse applications (e.g., speech recognition and object recognition in vision).
不在現場的我,是沒有機會親聆這場盛會啦,不過稍微瞄了下簡報檔,覺得這本書裡面的題材都蠻有意思的,比如說下面第一個圖提到克服隱私權疑慮的嘗試,和第二個圖 Tree Ensembles 的說明。
期待這本書的上市,不過看到價格實在是讓人有點遲疑啊,哈哈!
Sunday, April 10, 2011
失之毫釐是不是謬以千里
很多人使用 Likert Scale 做評分(Ratings)的量表基礎,比如說像「非常不喜歡、喜歡、無所謂、不喜歡、非常不喜歡」這樣的評分表就極爲常見,但是 Xavier 提醒我們 Likert Scale 的數據是 ordinal data ,這種數據僅僅表達次序關係,但是兩兩評分之間未必是 equidistant 的。若用這樣的數據計算距離(計算距離是相似性的基礎),其結果可能是失真的,循此邏輯推演下去,計算推薦系統準確率的指標 RMSE 的意義也可能失準。
從數學的角度來看,誤用定義當然是極爲嚴重的基本功的失誤,但是若從實務上考量,把 Likert 式評分當做 internal data,對推薦系統的成果究竟影響又多大,實在不好說 。不過,看來在這一點上不察,誤把馮京當馬涼的研究人員和開發人員可能不少哦!
Xavier Amatriain 寫這篇文章,是受 Judy Robertson 在 Blog@ACM 上的文章 We're Doing It Wrong 所啓發。Judy 在文中提到 2010 ACM Conference on Human Factors in Computing Systems 有學者發表研究 前一年會議中發表論文《Powerful and consistent analysis of Likert-type rating scales 》,爬梳學者使用的數據和統計工具,發現驚人的事實,原文是這樣的:
Kaptein, Nass, & Markopoulos (2010) published a paper in CHI last year found that in the previous year's CHI proceedings, 45% of the papers reported on likert type data but only 8% used non-parametric stats to do the analysis. 95% reported on small sample sizes (under 50 people). This is statistically problematic even if it gets past reviewers!
使用 Likert Scale 作爲實驗分析方法的學者竟然約略達到五成,Judy 在文章下半部提出她對此現象原因的觀察和建議,我對統計是大外行,只能點頭諾諾。但最抓住我眼球的句子是“95% reports on small sample size”這句,產業界鮮少有人信服學界真能做出「有用」的東西,確實有點道理,怨不得人。
[參考資料]
Kaptein, M., Nass, C., Markopoulos, P. (2010) Powerful and consistent analysis of Likert-type rating scales. In Proceedings CHI 2010, ACM, New York, NY, 2391-2394. DOI= http://doi.acm.org/10.1145/1753326.1753686
Sunday, January 30, 2011
How recommender researchers test their algorithm and make the system smarter?
There are many holy grails in online commerce, but one that has frustrated C-level executives and engineers alike is how to produce better recommendation algorithms. Produce better recommendations, and you’ll sell more stuff.
Historically, however, there has been one major structural impediment to making significant breakthroughs on this front. But the Chief Scientist at RichRelevance, which provides personalization solutions for the likes of Walmart, Sears, and Overstock.com, may just have fixed that.
First, however, the impediment: The people who are likely to produce breakthroughs--the really smart smarty-pants in the math departments of the world’s universities--don’t have access to large bodies of real-world data. And without real-world data, they can come up with as many hypotheses and new types of math as they like, but they’ll never really know if it actually works in the real world. It’s like trying to learn how to serve without tennis balls. You can swing as much as you like, but until you actually hit a real-live ball, you can never be sure if your swing would actually place a ball in the serve box.
For their part, the people who have real-world data--the Amazons and eBays of the world--can’t share it with the researchers for reasons of customer privacy. “Even if we anonymize it, we’re handcuffed because we can’t give out data that can be reasonably be used to reconstruct who someone really is,” the Chief Scientist, Darren Vengroff, tells Fast Company.
Vengroff, however, has come up with a novel solution: He’s created a “black box” of sorts with real-world data that researchers can use to run experiments on. Researchers won’t be able to look at the data, but they will be able to dump their algorithms in and have the box spit out results, which the researchers can then use to refine their hypotheses.
It’s a simple idea, but it wasn’t really possible to execute until the advent of the cloud. Now researchers from any part of the globe will be able to use the system to run experiments. (In principle, of course--in practice, a committee will vet proposals and choose which ones will actually run.)
Vengroff, who once worked as a Principal Engineer at Amazon, says he got the idea for the project while attending a computing conference last fall. “In one of the sessions, there were three consecutive papers in a row where about two-thirds of the way through, I was really excited about what was being presented, and then they went down a different path than I thought they were going to go,” he says. “I realized if they only knew what the real-live data, that I look at every day, says, they wouldn’t have gone that way. They would have gone the right way and gotten to a much better solution.”
“Seeing these brilliant ideas get misapplied because of a very reasonable assumption about how shoppers might shop, but happens not to be true in the empirical data--I realized we’ve got to find a way to have this not happen anymore.”
Monday, August 30, 2010
比開發推薦系統更好的賺錢方法
依照 Wired 的報導,其中兩個專利和推薦系統有關,一個專利是根據消費者這正在閱覽的頁面來決定向用戶推薦物品,另一個則是根據讀者正在閱讀文章的內容和超連接,向讀者推薦新聞故事。
Forbes 的專欄作者 Lee Gomes 對這則新聞的反應很直接:
This is yet another example of the cynical use of the American legal system to extort money out of successful companies — in the name of protecting innovation and innovators. Shame on Paul Allen for being part of it.
顯然 Paul Allen 瞄準的對象口袋深度都很夠,這齣戲的劇本很簡單,一個有錢人,伸手到其他有錢人的口袋撈錢。我不知道保羅大哥,除了賺更多錢之外還有什麼其他不可告人的深意,層次不到,無法揣測超級有錢人的心意。不過這時候回頭看看 Wired 1999 年底的文章Think Tanked,滋味很特別。
Resys 諸君,有沒有人想去念個法律學位呢!?
Saturday, March 27, 2010
Daniel Lemier's advice on How to Write Good Papers
Monday, March 8, 2010
Resys China 電子雜誌創刊號面世了
Tuesday, February 2, 2010
FANS AND EXPERTS OF RECOMMENDER SYSTEM IN TWITTER (by Geeksensor)
Monday, October 26, 2009
Francisco Martin #RecSys09 Industry Keynote Summary
Lesson 1 – Make sure a recommender is really needed! Do you have lots of recommendable items? Many diverse customers?… also think Return-on-Investment… a more sophisticated recommender may not deliver a better ROI.
Lesson 2 – Make sure the recommendations make strategic sense. Is the best recommendation for the customer also the best for the business? What is the difference between a good and useful recommendation? Good recommendations .vs. useful recs; Obvious recommendations may not be useful; risky recs may deliver better long-term value (所有系統都是為企業需求而生,切記切記)
Lesson 3 - Choose the right partner! Select the right rec vendor vs hire some #recsys09 students. If you are a big company the best you can do is to organize a contest (為什麼不直接明說 Netflix ?LOL)
Lesson 4 – Forget about cold-start problems (!) …. just be creative. The internet has the data you need (somewhere…) (記住那句老話:We are limited only by our imigination)
Lesson 5 – Get the right balance between data and algorithms. 70% of the success of a #recsys is on the data, the other 30% on the algorithm (這個問題我們已經討論很多次了, Worry about the data before you worry about the algorithm)
Lesson 6 – Finding correlated items is easy but deciding what, how, and when to present to the user is hard… or don't just recommend for the sake of it. Remember user attention is a scarce and valuable resource. Use it wisely! … don't make a recommendations to a customer who is just about to pay for items at the checkout! User interface should get at least 50% of your attention.
Lesson 7 – Don't waste time computing nearest neighbours (use social connections)… just mine the social graph. Might miss useful connections??
Lesson 8 – Don't wait to scale (6, 7, 8, 9 顯然都是實務上的經驗談)
Lesson 9 – Choose the right feedback mechanism. Stars vs thumbs …. the YouTube problem. More research on implicit and other feedback mechanisms is needed. The perfect rating system is no rating system! … focus on the interface. Seems to me this is one of the gaps in current research… algorithms > data > interface
Lesson 10 – Measure Everything! … business control and analytics is a big opportunity here. (不僅要評量預測準不準,企業流程裡每個環節都要有評估機制,這是有真正創業、經營體驗的人的心得)
Keynote Takeaway – Think about application context; Focus on interface as much as algorithms; Be creative with start-up data. … the UI needs to get the lion’s share of the effort (50%) compared to algorithms (5%) , knowledge (20%), analytics (25%)
對於最後的 Takeaway,每個讀者或許都有自己的看法,畢竟要量化各因素在系統開發過程中的比重實在不容易,最後只能是被迫給出一組表達自己“經驗值”的數字。UI 的重要性當然毋庸置疑,只是 UI 為什麼是演算法的十倍?聰明的你(妳),想必有一套自己的想法!
Sunday, October 25, 2009
#RecSys09 話題: what's on recommender researchers' mind?
如果想更快知道究竟今年有哪些熱門話題,就來看看在 University College Dublin 教書的 Barry Smyth 為大家製作的標籤雲,看起來 Netflix 還是大熱門的 buzzword 啊!
Sunday, October 4, 2009
關於 user-based 和 item-based 算法的思考
半夜被地震閙醒,起床上網瞎晃,看到十一假期間xlvector兄仍然勤奮不輟思考 user-based 和 item-based 演算法對於輸出多様化的比較,大為佩服,特記錄于後:
(此刻鄙人凖備爬上床睡回龍覺)
前一段时间和wendong聊天,他提到userbased算法的结果多样性不如itembased算法。对此,我觉得有几个问题
1) 我们知道所谓多样性,是指推荐结果两两都不怎么相似,从而不同的相似度度量其实产生不同的多样性度量。
2)常用的相似度有两种,一种是基于content的,一种是基于collaborative filtering的,那么根据我的实验,在这两种相似度的度量下,userbased的结果多样性都好于itembased的算法
3)但我觉得还是存在一种相似度,而这个相似度对应的多样性在item-based的方法下比较好
不知道大家对这个问题怎么看,在实际系统中userbased和itembased谁能产生多样的结果?
Wednesday, September 30, 2009
關於 Noise 的隨想與牢騷

Tuesday, August 25, 2009
豆辦解決圖書版本問題的方法
這個現象解釋起來很簡單,很多產品有不同的包裝版本(想像一下書籍的普及版、精裝本、典藏版, 等等),甚至同樣的版本,也可能在資料庫裡有兩筆以上資料(仍然以書籍做例子,想像一下同一本書的一刷、二刷),我們如何知道這些不同的商品其實都是同一個產品?Greg Linden 也曾經以 YouTube 為例,撰文 YouTube cries out for item authority 說明 item authority問題對服務提供者造成的困擾以及挑戰。
這對推薦系統的設計者,是個很「有趣」,也很艱鉅的挑戰!海峽兩岸間最大的書籍收藏網站-豆辦,當然也遇到這個問題,豆辦日誌昨日(2009/08/24)宣佈《豆瓣读书即将解决版本问题》,他們的解決方案,可以稱為用戶「自己動手,豐衣足食」吧:
豆瓣图书会将同一作品的不同版本归纳起来,展示在一个单独页面里。这个页面可以由书虫们来添加和编辑。如果你确切地知道06年上海译文出版社的《在路上》是01年漓江出版社的《在路上》的另一个版本,你可以添加;如果发现某个版本是指鹿为马,你可以报错。贡献者的信息会在版本页面被永久标记。
随着豆瓣数据库里的版本数据的完善,豆瓣猜的智商也将大大提高,再也不会推荐同一作品不同版本的书给你了;有些已经绝版不再出售的图书页面(比如 86年版的《傲慢与偏见》),会有最近新版的价格帮助购买(比如06年的《傲慢与偏见》有售);对于多达十几种版本的图书,版本页面还会显示各自的收藏人数和评分,帮助大家比较版本的好坏。
Thursday, August 20, 2009
推薦系統不是只能賣書
因為我找到的 PDF 連結是壞的,還沒有機會看到全文,所以只能從摘要想像大概。不過,這的確是個有趣的想法,推薦系統不僅在 attention economy 大環境裡協助企業主攫取客戶注意力,還能協助程式設計師更好的完成他的工作(assignment),看來 "We're limited only by our imagination" 這句話真是一點都沒錯。
DebugAdvisor: A Recommender System for Debugging
In large software development projects, when a programmer is assigned a bug to fix, she typically spends a lot of time searching (in an ad-hoc manner) for instances from the past where similar bugs have been debugged, analyzed and resolved. Systematic search tools that allow the programmer to express the context of the current bug, and search through diverse data repositories associated with large projects can greatly improve the productivity of debugging. This paper presents the design, implementation and experience from such a search tool called DebugAdvisor.
The context of a bug includes all the information a programmer has about the bug, including natural language text, textual rendering of core dumps, debugger output etc.
Our key insight is to allow the programmer to collate this entire context as a query to search for related information. Thus, DebugAdvisor allows the programmer to search using a fat query, which could be kilobytes of structured and unstructured data describing the contextual information for the current bug. Information retrieval in the presence of fat queries and variegated data repositories, all of which contain a mix of structured and unstructured data is a challenging problem. We present novel ideas to solve this problem.
We have deployed DebugAdvisor to over 100 users inside Microsoft. In addition to standard metrics such as precision and recall, we present extensive qualitative and quantitative feedback from our users.
Monday, July 27, 2009
Er....Netflix Prize goes to...?
原本以為最終結果應該是 BPC 的囊中物,應該不會有什麼懸念,但就在時間截止之前,另外一個隊伍 Ensemble 宣稱他們也跨過門檻(Breaking - Netflix Prize, we’ve got a winner, and it’s Greek! (updated)),甚至一度自行宣佈他們勝過原本的領先者,是最終的贏家(下圖是目前Leaderboard公佈的成績)。不過,一位自稱 An Insider 的網友在這篇文章之後留言,解釋他們可能誤解了規則,BPC 在 Test Set 的成績較優,才是最終的贏家(自稱 Just a guy in garage 的 Gavin Potter 很快的在個人部落格撰文解釋為什麼 BPC 才是贏家)。
到目前為止, Netflix 仍然沒有正式宣佈誰是最後的贏家,只是宣佈停止收件,並且說有兩個隊伍通過門檻。
As of July 26, 2009 18:42:37 UTC, we have stopped gathering submissions for the Netflix Prize contest. There are submissions from two teams that meet the minimum requirements for the Grand Prize. We are contacting the lead team and we will report, as soon as possible, when and if we have a verified winner for the Grand Prize.
補充:
不管誰贏得這個比賽, Daniel Lemire 說的好,充份鼓勵各種創意和多元化的發展,才是學術發展的正確方向:
Both teams broke the 10% barrier by using a diverse coalition, by merging several different ideas. As Peter Turney recently stated:I am now more convinced than ever that science needs diverse explanations, techniques and opinions. We should actively reward creativity. Science is not merely about truth-seeking.
There are no whole-truths, but we can get by reasonably well with a large number of half-truths.
Saturday, June 27, 2009
Netflix Prize goes to BellKor's Pragmatic Chaos?
BellKor's Pragmatic Chaos 在他們的網頁上宣佈,他們已經達到 10.05% 的成果,所以他們贏得本次大獎賽。
June 26, 2009: Today our team submitted our solution to the Netflix Prize, resulting in a score of .8558, which corresponds to an improvement over Netflix Cinematch algorithm of 10.05%. This is the first submission in the competition to break the 10% barrier and sets off a 30 day period where all competitors are invited to submit their best and final solutions.
此刻 Netflix Prize Leaderboard 也顯示 BPC 在 2009-06-26 18:42:37 提出的成果已經達到 10% 門檻值:
Update 1:
RWW 隨後也報導了這個消息, Marshall Kirkpatrick 執筆的 They Did It! One Team Reports Success in the $1m Netflix Prize 文中簡述比賽的背景,並且引用紐約時報的專文 If You Liked This, You’re Sure to Love That 談到所謂正確預測用戶的喜好,不見得是萬靈丹。下一代的推薦系統,還有很長的路要走。
Update 2:
真正的推薦系統行家人Greg Linden 顯然知道更多內情,他不僅說明 BPC 是由四個團隊合併合成,還提醒大家一件事,別的團隊有一個月的緩衝時間逆轉局勢:Other teams have 30 days to beat it, but, no matter what happens, the $1M prize will be claimed in the next month。
不論最後結局如何,恭喜得勝的團隊,Congratulations to the winners!
Friday, January 23, 2009
Question: Netflix Prize within months?
雅虎的資深研究員 John Langford 在閱讀 BellKor in Chaos 釋出的文件(1,2)後,在他的個人部落格 Machine Learning (Theory) 指出也許距離獎金揭曉的時間不遠了,他同時指出該團隊演算法包含了 stochastic gradient descent, ensemble prediction, and targeting residuals 各領域的技術,並且指出他們在 2008 年間將演算法參數化的努力。他同時意味深長的說,或許- the right parameterization might very well succeed - 正確的參數就能將大獎抱走呢!
Several aspects of solutions are taken for granted including stochastic gradient descent, ensemble prediction, and targeting residuals (a form of boosting). Relatively to last year, it appears that many approaches have added parameterizations, especially for the purpose of modeling through time.當然,John Langford 也在他的短文裡,提出了他的疑慮:One fear is that the progress is asymptoting on the wrong side of the 10% threshold,可是他也承認從去年底到今年一月的進步的的確令人印象深刻 。
不過我個人比較好奇的是這些演算法,能不能作適當的改變後,應用到其他的產品,或許參數化正是往這個方向努力的指標之一,但是我對 overfitting 這件事仍然有點疑慮。雖然還沒有時間深入研讀相關文件,但在 John Langford 的部落格留言區以及 Netflix 網站的論壇裡有讀者討論到 overfitting 的問題,雖然有人認為這個比賽已經在各方面取得平衡(請參考這裡),不需擔心 overfitting,但是我個人仍然存疑。
或許 - Time will tell。
Wednesday, December 17, 2008
Netflix Progress Prize for 2008 宣布了
It is our great honor to announce the winner of the Netflix Progress Prize for 2008 as team BellKor in BigChaos for their verified just-in-time submission on Sept 30 at 21:17:40 UTC achieving a 9.44% improvement over Cinematch. We congratulate the team of Yehuda Koren, Robert Bell and Chris Volinsky of AT&T Research Labs combined with Andreas Töscher and Michael Jahrer of Commendo Research for their superb work integrating many significant techniques to achieve this result.
In accord with the Rules the team has prepared a system description consisting of two papers, which we both make public below. We will be awarding the Prize in a presentation at the Netflix offices in Los Gatos on December 17, 2008 at 4pm. Andreas Töscher and Michael Jahrer will present a public talk at that time about their Prize algorithm. We will post a video of that presentation via the Forum.
BellKor 團隊在網站上提供該團隊所發表與本次競賽有關的論文,供有興趣的讀者下載參考:
Tuesday, December 16, 2008
Reading List: Diversity in Recommenders
[1] C. Clarke, M. Kolla, G. Cormack, O. Vechtomova, A. Ashkan, S. Büttcher and I. Mackinnon, "Novelty and diversity in information retrieval evaluation," in SIGIR '08: Proceedings of the 31st Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, 2008, pp. 659-666.
[2] D. Fleder and K. Hosanagar, "Blockbuster Culture's Next Rise or Fall: The Impact of Recommender Systems on Sales Diversity," SSRN eLibrary, 2008.
[3] D. Fleder and K. Hosanagar, "Recommender systems and their impact on sales diversity," in EC '07: Proceedings of the 8th ACM Conference on Electronic Commerce, 2007, pp. 192-199.
[4] L. Iaquinta, M. de Gemmis, P. Lops, G. Semeraro, M. Filannino and P. Molino, "Introducing Serendipity in a Content-Based Recommender System," Hybrid Intelligent Systems, 2008. HIS '08. Eighth International Conference on, pp. 168-173, 2008.
[5] Q. Le and A. Smola, "Direct Optimization of Ranking Measures," Apr 2007. [Online]. Available: http://arxiv.org/abs/0704.3359.
[6] D. Lemire, S. Downes and S. Paquet, "Diversity in open social networks," 2008.
[7] L. Mcginty and B. Smyth, "On the Role of Diversity in Conversational Recommender Systems," 2003.
[8] S. Mcnee, J. Riedl and J. Konstan, "Being accurate is not enough: How accuracy metrics have hurt recommender systems," in CHI '06: CHI '06 Extended Abstracts on Human Factors in Computing Systems, 2006, pp. 1097-1101.
[9] K. Swearingen and R. Sinha, "Beyond algorithms: An HCI perspective on recommender systems," 2001.
[10] Y. Xu and H. Yin, "Novelty and topicality in interactive information retrieval," J. Am. Soc. Inf. Sci. Technol., vol. 59, pp. 201-215, 2008.
[11] C. Zhai, W. Cohen and J. Lafferty, "Beyond independent relevance: Methods and evaluation metrics for subtopic retrieval," in SIGIR '03: Proceedings of the 26th Annual International ACM SIGIR Conference on Research and Development in Informaion Retrieval, 2003, pp. 10-17.
[12] F. Zhang, "Research on Recommendation List Diversity of Recommender Systems," Management of e-Commerce and e-Government, 2008. ICMECG '08. International Conference on, pp. 72-76, 2008.
[13] M. Zhang and N. Hurley, "Avoiding monotony: Improving the diversity of recommendation lists," in RecSys '08: Proceedings of the 2008 ACM Conference on Recommender Systems, 2008, pp. 123-130.
Tuesday, July 22, 2008
協同過濾(collaborative filtering)推薦系統的實作
最近讀了交通大學資管所劉敦仁教授2007年發表在 Expert Systems with Application 的文章[1] ,他將客戶終生價值(customer lifetime value, CLV)融入協同過濾(collaborative filtering,CF)推薦系統框架,以加權後的 RFM (Recency、Frequency、Monetary)模型,作為客戶分群的依據。
試著將更多實務界或商管領域的思維,整合至資料挖掘的實作,一直是我在思考的方向,這篇文章的思路對我並不陌生,因此我試著更深入理解他的做法。
在閱讀的過程裡,我覺得這篇文章在整理過去研究成果(related work)的部分,蠻有意思的,本文整理協同過濾的各種不同做法,以實作時應用的各種基礎演算法(例如:關聯法則、分群)為基礎的分類思路,而不是從商品與顧客的不同觀察視角(item-based .vs. user-based)出發。個人認為,這種切入角度,能夠幫助有意實作推薦系統的讀者,更快的理解推薦系統的組成架構,並且幫助他們更有效率的擬定工作計畫。
協同過濾的基本精神,在於數大就是美,資料愈多,系統的表現愈佳。協同過濾的實作精要之處,則在於如何從購物人潮中找出與特定顧客品味嗜好相近的同好,或是任意揀選一件商品,如何找出相似的品項。如何找出人與物的相似處,就取決於相似度(similarity)的計算方式了。歷來學者曾在文獻中建議使用的相似度公式,五花八門琳琅滿目,用族繁不及備載來形容一點也不誇張。
最常被人提及的計算方式包括 Euclidean Distance、 Pearson correlation coefficient、Jaccard coefficient、Manhatten distance、Cosine correlation coefficient 等等。許多學者在這些基礎上,設計了更複雜的計算方式,比如劉教授建議以商管領域常使用的 RFM (Recency、Frequency、Monetary)模型,計算客戶貢獻度(客戶對業者的價值)為基礎的計算方式,他還以此為基礎,建議更複雜的加權式 RFM (Weighted RFM)計算公式。簡而言之,更有用的相似度判斷方式,一直是學者努力的重點之一。
定義相似度之後,最重要的是怎麼應用相似度來建構推薦機制。根據劉教授的整理,有三大類計算方式(當然啦,這只是我個人的理解):
k-nearest neighbor(kNN )
這個方法是最直覺,也最容易理解的。指定一個消費者(或者指定商品,道理都是一樣的),利用相似度公式,計算出和這個消費者最相似的 k 個顧客。然後我們分析這些選出來的對象,找出他們購買過,但我們的主人翁還沒有購買的項目,這些就是要推薦給主人翁的商品。實務上,可能在計算上更複雜一點,不過這就是最基本的道理了。如果讀者可以參考 Programming Collective Intelligence (O'REILLY 出版)這本書的第二章和第八章,就更能體會這個方法的神髓了。
這個方法在邏輯上也很單純,先將顧客依照某些條件分群,讓後再針對每一個群組作關聯法則的分析。我們依照關聯法則的分析,對這個群組裡的消費者,產生推薦清單(這個群裡其他人都買了,他還沒有買的商品就是推薦對象)。最簡單直覺的分群演算法,就是 k-mean clustering 了,但別忘了最重要的分群依據,就是前述一再強調的相似度計算方法。
寸有所長,尺有所短。每一種方法,都有其優勢和缺點,因此學者嘗試將 content-based (CB) 和 collaborative filtering (CF)結合在一起,這就是所謂的融合解法 (hybrid approach)。將不同的計算方法結合在一起,說來簡單,作起來卻有很多變化。學者 Burke 有篇論文[2] - Hybrid Recommender Systems: Survey and Experiments,整理了做法,一共有 weighted、switching、mixed、feature combination、casade、feature augmentaion & meta-level 七種之多。依照劉教授的整理看來,他最重視 weighted 和 meta-level 兩種方法。
雖然我一直覺得,RFM 是不是能表達愛好與品味,還是個問號?不過本篇論文的整理功夫,的確值得稱道,配合 Programming Collective Intelligence 一起看,將論文裡的數學符號和實作連結起來,收獲特別多。
參考資料:
[1] Y. Shih and D. Liu, “Product recommendation approaches: Collaborative filtering via customer lifetime value and customer demands,” Expert Syst. Appl., vol. 35, 2008, pp. 350-360.
[2] R. Burke, “Hybrid Recommender Systems: Survey and Experiments,” User Modeling and User-Adapted Interaction, vol. 12, Nov. 2002, pp. 331-370.
如果我的心是一朵蓮花
~ 林徽因 · 馬雁散文集 · 蓮燈 ~ 馬雁 在她的散文《高貴一種,有詩為證》裡,提到「十多年前,還不知道林女士的八卦及成就前,在期刊上讀到別人引用的《蓮燈》」 覺得非常喜歡,比之卞之琳、徐志摩,別說是毫不遜色,簡直是勝出一籌。前面的韻腳和平仄的處理顯然高於戴...
-
~ 林徽因 · 馬雁散文集 · 蓮燈 ~ 馬雁 在她的散文《高貴一種,有詩為證》裡,提到「十多年前,還不知道林女士的八卦及成就前,在期刊上讀到別人引用的《蓮燈》」 覺得非常喜歡,比之卞之琳、徐志摩,別說是毫不遜色,簡直是勝出一籌。前面的韻腳和平仄的處理顯然高於戴...
-
在資通訊安全領域,有個源於軍事領域的術語 defense in depth (DID), 在軍事上,DID 是一種 策略 理念,防禦不能只靠一道強大的防線(比如說中國的長城和法國的 Maginot Line ),必須用多層次(multi-layer)、多角度、多點、多面的防...

