Alon Halevy, Peter Norvig, and I argue that we should stop acting as if our goal is to author extremely elegant theories, and instead embrace complexity and make use of the best ally we have: the unreasonable effectiveness of data. See the full article here (IEEE Intelligent Systems, March/April 2009).
Wednesday, 25 March 2009
The Unreasonable Effectiveness of Data
Posted on 17:02 by Unknown
Alon Halevy, Peter Norvig, and I argue that we should stop acting as if our goal is to author extremely elegant theories, and instead embrace complexity and make use of the best ally we have: the unreasonable effectiveness of data. See the full article here (IEEE Intelligent Systems, March/April 2009).
Thursday, 19 March 2009
Google and WPP Marketing Research Awards: Improving industry understanding and practices in online marketing
Posted on 11:10 by Unknown
Posted by Jeff Walz, University Relations and Anne Bray, Head of Agency, WPP
Google and the WPP Group have teamed up to create a new research program with the goal of improving industry understanding of digital marketing. Eleven research awards have been given to universities through the Google and WPP Marketing Research Awards Program, announced by both companies in the fall of 2008. The academic studies will harness WPP client data to explore how online media influences consumer behavior, attitudes, and decision making. The research provides an opportunity for very innovative thinking in an area that is at the crossroads of marketing, computer science, economics, and various mathematical disciplines.
More than 120 entries were received by the deadline for proposals. The awards represent the first round of grants in the three-year program towards which WPP and Google will commit up to $4.6 million in an effort to support research around digital marketing. Hal Varian, Google's Chief Economist, participated on the decision committee.
The winning projects offer convincing designs for exploring how online and offline marketing influence consumer attitudes, decisions, and purchase behavior. As marketing continues to become more digital and more measurable, the results of these studies will also advance our understanding of how advertising investment should be allocated among media channels.
The researchers and affiliated academic institutions participating in this first round of awarded projects are:
• “Effect of Online Exposure on Offline Buying: How Online Exposure
Aids or Hurts Offline Buying by Increasing the Impact of Offline
Attributes”; Amitav Chakravarti, New York University, Stern School of
Business, Department of Marketing
• “The Interaction Between Digital Marketing Tactics and Sales
Performance Online and Offline”; Elie Ofek, Associate Professor
Marketing, Harvard Business School and Zsolt Katona, Associate
Professor of Marketing, UC Berkeley, Haas School of Business
• ”Are Brand Attitudes Contagious? Consumer Response to Organic
Search Trends”; Donna L. Hoffman, Professor, A. Gary Anderson
Graduate School of Management, University of California Riverside and
Thomas P. Novak, A. Gary Anderson Graduate School of Management,
University of California Riverside
• “Does internet advertising help established brands or niche ("long
tail") brands more? Catherine Tucker, Assistant Professor of
Marketing, MIT Sloan School of Marketing and Avi Goldfarb, Associate
Professor of Marketing, Joseph L. Rotman School of Management
University of Toronto
• “Marketing on the Map: Visual Search and Consumer Decision Making”;
Nicolas Lurie, Assistant Professor of Marketing, College of
Management, Georgia Institute of Technology, College of Management and
Sam Ransbotham, Assistant Professor of Information Systems, Carroll
School of Management, Boston College
• “Methods for multivariate metric analysis; identifying change
drivers”; Trevor J. Hastie, Professor, Department of Statistics,
Stanford University
• “Unpuzzling the Synergy of Display and Search Advertising: Insights
from Data Mining of Chinese Internet Users”; Hairong Li, Department of
Advertising, Public Relations, and Retailing, Michigan State
University and Shuguang Zhao, Media Survey Lab, Tsinghua University
• “Optimal Allocation of Offline and Online Media Budget”; Sunil
Gupta, Professor of Business Administration, Harvard Business School;
Anita Elberse, Associate Professor, Harvard Business School; and
Kenneth C. Wilbur, Assistant Professor of Marketing, Marshall School
of Business, University of Southern California
• “Targeting Ads to Match Individual Cognitive Styles: A Market
Test”; Glen Urban, Professor, MIT Sloan School of Management
• “How do consumers determine what is relevant? A psychometric and
neuroscientific study of online search and advertising effectiveness”;
Antoine Bechara, Professor of Psychology and Neuroscience, Department
of Psychology/Brain & Creativity Institute, University of Southern
California and Martin Reimann, Fellow, Department of Psychology/Brain
& Creativity, University of Southern California
• “A Comprehensive Model of the Effects of Brand-Generated and
Consumer-Generated Communications on Brand Perceptions, Sales and
Share”; Douglas Bowman and Manish Tripathi, Professors of Marketing,
Goizueta Business School, Emory University.
You can find more information about the Google and WPP Marketing Research Awards Program on the website.
Google and the WPP Group have teamed up to create a new research program with the goal of improving industry understanding of digital marketing. Eleven research awards have been given to universities through the Google and WPP Marketing Research Awards Program, announced by both companies in the fall of 2008. The academic studies will harness WPP client data to explore how online media influences consumer behavior, attitudes, and decision making. The research provides an opportunity for very innovative thinking in an area that is at the crossroads of marketing, computer science, economics, and various mathematical disciplines.
More than 120 entries were received by the deadline for proposals. The awards represent the first round of grants in the three-year program towards which WPP and Google will commit up to $4.6 million in an effort to support research around digital marketing. Hal Varian, Google's Chief Economist, participated on the decision committee.
The winning projects offer convincing designs for exploring how online and offline marketing influence consumer attitudes, decisions, and purchase behavior. As marketing continues to become more digital and more measurable, the results of these studies will also advance our understanding of how advertising investment should be allocated among media channels.
The researchers and affiliated academic institutions participating in this first round of awarded projects are:
• “Effect of Online Exposure on Offline Buying: How Online Exposure
Aids or Hurts Offline Buying by Increasing the Impact of Offline
Attributes”; Amitav Chakravarti, New York University, Stern School of
Business, Department of Marketing
• “The Interaction Between Digital Marketing Tactics and Sales
Performance Online and Offline”; Elie Ofek, Associate Professor
Marketing, Harvard Business School and Zsolt Katona, Associate
Professor of Marketing, UC Berkeley, Haas School of Business
• ”Are Brand Attitudes Contagious? Consumer Response to Organic
Search Trends”; Donna L. Hoffman, Professor, A. Gary Anderson
Graduate School of Management, University of California Riverside and
Thomas P. Novak, A. Gary Anderson Graduate School of Management,
University of California Riverside
• “Does internet advertising help established brands or niche ("long
tail") brands more? Catherine Tucker, Assistant Professor of
Marketing, MIT Sloan School of Marketing and Avi Goldfarb, Associate
Professor of Marketing, Joseph L. Rotman School of Management
University of Toronto
• “Marketing on the Map: Visual Search and Consumer Decision Making”;
Nicolas Lurie, Assistant Professor of Marketing, College of
Management, Georgia Institute of Technology, College of Management and
Sam Ransbotham, Assistant Professor of Information Systems, Carroll
School of Management, Boston College
• “Methods for multivariate metric analysis; identifying change
drivers”; Trevor J. Hastie, Professor, Department of Statistics,
Stanford University
• “Unpuzzling the Synergy of Display and Search Advertising: Insights
from Data Mining of Chinese Internet Users”; Hairong Li, Department of
Advertising, Public Relations, and Retailing, Michigan State
University and Shuguang Zhao, Media Survey Lab, Tsinghua University
• “Optimal Allocation of Offline and Online Media Budget”; Sunil
Gupta, Professor of Business Administration, Harvard Business School;
Anita Elberse, Associate Professor, Harvard Business School; and
Kenneth C. Wilbur, Assistant Professor of Marketing, Marshall School
of Business, University of Southern California
• “Targeting Ads to Match Individual Cognitive Styles: A Market
Test”; Glen Urban, Professor, MIT Sloan School of Management
• “How do consumers determine what is relevant? A psychometric and
neuroscientific study of online search and advertising effectiveness”;
Antoine Bechara, Professor of Psychology and Neuroscience, Department
of Psychology/Brain & Creativity Institute, University of Southern
California and Martin Reimann, Fellow, Department of Psychology/Brain
& Creativity, University of Southern California
• “A Comprehensive Model of the Effects of Brand-Generated and
Consumer-Generated Communications on Brand Perceptions, Sales and
Share”; Douglas Bowman and Manish Tripathi, Professors of Marketing,
Goizueta Business School, Emory University.
You can find more information about the Google and WPP Marketing Research Awards Program on the website.
Wednesday, 18 March 2009
And the award goes to...
Posted on 09:22 by Unknown
Posted by Fernando Pereira, Research Director
Corinna Cortes, Head of Google Research in New York, has just been awarded the ACM Paris Kanellakis Theory and Practice Award jointly with Vladimir Vapnik (Royal Holloway College and NEC Research). The award recognizes their invention in the early 1990s of the soft-margin support vector machine, which has become the supervised machine learning method of choice for applications ranging from image analysis to document classification to bioinformatics.
What is so important about this invention? In supervised machine learning, we create algorithms that can learn a rule to accurately classify new examples based on a set of training examples (e.g. spam or non-spam). There is no single attribute of an email message that tells us with certainty that it is spam. Instead, many attributes have to be considered, forming a vector of very high dimension. The same situation arises in many other machine practical learning tasks, including many that we work on at Google.
To learn accurate classifiers, we need to solve several big problems. First, the rule learned from the training data should be accurate on new test examples, even though it has not seen those examples. In other words, the rule must generalize well. Second, we must be able to find the optimal rule efficiently. Both of these problems are especially daunting for very high dimensional data. Third, the method for computing the rule should be able to accommodate errors in the training data, such as messages that are given conflicting labels by different people (my spam may be your ham).
Soft-margin support vector machines wrap these three problems together into an elegant mathematical package. The crucial insight is that classification problems of this kind can be expressed as finding in very high dimension (or even infinite dimension) the hyperplane that best separates the positive examples (ham) from the negative ones (spam).
Remarkably, the solution of this problem does not depend on the dimensionality of the data, it depends only on the pairwise similarities between the training examples determined by the agreement or disagreement between corresponding attributes. Furthermore, a hyperplane that separates the training data well can be shown to generalize well to unseen data with the same statistical properties.
Now, you might be asking how could this be done if the training data is inconsistently labeled. After all, you cannot have the same example on both sides of the separating hyperplane. That's where the soft margin idea comes in: the quadratic optimization program that finds the optimal separating hyperplane can be cleverly modified to "give up" on a fraction of the training examples that cannot be classified correctly.
With this crucial improvement, support vector machines became really practical, while the core ideas have had huge influence in the development of further learning algorithms for an ever wider range of tasks.
Congratulations to Corinna (and Vladimir) on the well-deserved award.
Corinna Cortes, Head of Google Research in New York, has just been awarded the ACM Paris Kanellakis Theory and Practice Award jointly with Vladimir Vapnik (Royal Holloway College and NEC Research). The award recognizes their invention in the early 1990s of the soft-margin support vector machine, which has become the supervised machine learning method of choice for applications ranging from image analysis to document classification to bioinformatics.
What is so important about this invention? In supervised machine learning, we create algorithms that can learn a rule to accurately classify new examples based on a set of training examples (e.g. spam or non-spam). There is no single attribute of an email message that tells us with certainty that it is spam. Instead, many attributes have to be considered, forming a vector of very high dimension. The same situation arises in many other machine practical learning tasks, including many that we work on at Google.
To learn accurate classifiers, we need to solve several big problems. First, the rule learned from the training data should be accurate on new test examples, even though it has not seen those examples. In other words, the rule must generalize well. Second, we must be able to find the optimal rule efficiently. Both of these problems are especially daunting for very high dimensional data. Third, the method for computing the rule should be able to accommodate errors in the training data, such as messages that are given conflicting labels by different people (my spam may be your ham).
Soft-margin support vector machines wrap these three problems together into an elegant mathematical package. The crucial insight is that classification problems of this kind can be expressed as finding in very high dimension (or even infinite dimension) the hyperplane that best separates the positive examples (ham) from the negative ones (spam).
Remarkably, the solution of this problem does not depend on the dimensionality of the data, it depends only on the pairwise similarities between the training examples determined by the agreement or disagreement between corresponding attributes. Furthermore, a hyperplane that separates the training data well can be shown to generalize well to unseen data with the same statistical properties.
Now, you might be asking how could this be done if the training data is inconsistently labeled. After all, you cannot have the same example on both sides of the separating hyperplane. That's where the soft margin idea comes in: the quadratic optimization program that finds the optimal separating hyperplane can be cleverly modified to "give up" on a fraction of the training examples that cannot be classified correctly.
With this crucial improvement, support vector machines became really practical, while the core ideas have had huge influence in the development of further learning algorithms for an ever wider range of tasks.
Congratulations to Corinna (and Vladimir) on the well-deserved award.
Wednesday, 18 February 2009
Beyond Web-2.0
Posted on 16:08 by Unknown
Posted by T.V Raman, Research Scientist
A little over a year ago, I gave a lightning talk at the W3C Technical Plenary in Boston where I looked forward to what came after Web-2.0. The key insight underlying that talk was that the Web was now mature enough for us to build Web technologies purely out of Web parts. Web-2.0 is a result of applying the Web to itself and is therefore better thought of as Web(Web()) or more concisely, Web2.
Today, we can build new web artifacts out of existing ones by aggregation (web mashups) and by projection (filtered views), and publish the resulting artifacts on the web by assigning them a URL. This leads to the insight that this web that is to come potentially consists of the power-set of all web content. These ideas, and their logical consequences are detailed in article entitled Toward 2^W --- Beyond Web 2.0 in the February edition of the Communications Of The ACM. You can also find a slightly more extensive blog I posted here.
A little over a year ago, I gave a lightning talk at the W3C Technical Plenary in Boston where I looked forward to what came after Web-2.0. The key insight underlying that talk was that the Web was now mature enough for us to build Web technologies purely out of Web parts. Web-2.0 is a result of applying the Web to itself and is therefore better thought of as Web(Web()) or more concisely, Web2.
Today, we can build new web artifacts out of existing ones by aggregation (web mashups) and by projection (filtered views), and publish the resulting artifacts on the web by assigning them a URL. This leads to the insight that this web that is to come potentially consists of the power-set of all web content. These ideas, and their logical consequences are detailed in article entitled Toward 2^W --- Beyond Web 2.0 in the February edition of the Communications Of The ACM. You can also find a slightly more extensive blog I posted here.
Wednesday, 28 January 2009
Market Algorithms and Optimization Meeting
Posted on 10:25 by Unknown
Posted by Vahab S. Mirrokni and Muthu Muthukrishnan
Google auctions ads, and enables a market with millions of advertisers and users. This market presents a unique opportunity to test and refine economic principles as applied to a very large number of interacting, self-interested parties with a myriad of objectives. Researchers in economics, computer science, operations research, marketing and business are increasingly involved in defining, understanding and influencing this market.
On January 7th, a group of computer science researchers from Google and various universities formed a workshop to discuss key research directions in this area, specifically "Market Algorithms and Optimization." Corinna Cortes (Head, Google Research, NY) and Alfred Spector (VP of Research, Google) gave short talks about research groups in Google, academic collaborations, and research awards. The morning session comprised talks by Google researchers including Noam Nisan, Jon Feldman, Vahab Mirrokni, Yishay Mansour, and Muthu Muthukrishnan. These talks explored the following topics and key issues:
The post-lunch (sushi included of course) session comprised talks by researchers from academia. Bobby Kleinberg (Cornell) went first and took the audience far into projective geometry in an attempt to understand equilibria resulting from a sequence of selfish behavior of players using specific learning algorithms. Silvio Micali (MIT) described how difficult it was to model or propose mechanisms in presence of collusions, and then went on to show how to correctly design such mechanisms for combinatorial auctions. Kamesh Munagala (Duke Univ.) and Anna Karlin ( U. Wash) discussed algorithmic problems related to ad auctions, including how to run auctions in presence of advertisers with mixed utilities and tight remnant budgets. Other participants -- Richard Cole (NYU), Amos Fiat (Tel Aviv), Uri Feige (Weizmann), Michel Goemans (MIT), Anupam Gupta (CMU), Nicole Immorlica (Northwestern), Michael Rabin (Harvard), and Eva Tardos (Cornell) -- kept it lively with questions and discussions.
Part of what makes Google ad systems exciting is the challenges they pose and overcome, and yet others in the horizon. It is remarkable how some of the fundamental problems Google ad systems grapple with are also some of the hardest research problems in the community of Market Algorithms. Joint Google and Academia meetings like this help researchers begin to attack these problems, and may be a model for research collaborations in the future.
Google auctions ads, and enables a market with millions of advertisers and users. This market presents a unique opportunity to test and refine economic principles as applied to a very large number of interacting, self-interested parties with a myriad of objectives. Researchers in economics, computer science, operations research, marketing and business are increasingly involved in defining, understanding and influencing this market.
On January 7th, a group of computer science researchers from Google and various universities formed a workshop to discuss key research directions in this area, specifically "Market Algorithms and Optimization." Corinna Cortes (Head, Google Research, NY) and Alfred Spector (VP of Research, Google) gave short talks about research groups in Google, academic collaborations, and research awards. The morning session comprised talks by Google researchers including Noam Nisan, Jon Feldman, Vahab Mirrokni, Yishay Mansour, and Muthu Muthukrishnan. These talks explored the following topics and key issues:
- Role of budgets. Identify suitable properties of mechanisms for repeated auctions in the presence of budgets. Truthfulness is impossible, and may not be the suitable property.
- Advertisers' bidding strategies. While bidding equally on all keywords has certain desirable properties, identify other bidding strategies if advertisers want to maximize different utility functions and use rich bidding features.
- Dynamics of ad games. Understand dynamics of sophisticated advertiser bidder strategies under various ad allocation rules. For proportional allocation rules, the dynamics of simple bidder strategies are described here.
- Online Reservations. Design suitable mechanisms for selling ad inventory in the future rather than on the spot. Models and methods for what happens when reservations can not be honored are explored in WWW08, WINE08, and SODA09 papers.
The post-lunch (sushi included of course) session comprised talks by researchers from academia. Bobby Kleinberg (Cornell) went first and took the audience far into projective geometry in an attempt to understand equilibria resulting from a sequence of selfish behavior of players using specific learning algorithms. Silvio Micali (MIT) described how difficult it was to model or propose mechanisms in presence of collusions, and then went on to show how to correctly design such mechanisms for combinatorial auctions. Kamesh Munagala (Duke Univ.) and Anna Karlin ( U. Wash) discussed algorithmic problems related to ad auctions, including how to run auctions in presence of advertisers with mixed utilities and tight remnant budgets. Other participants -- Richard Cole (NYU), Amos Fiat (Tel Aviv), Uri Feige (Weizmann), Michel Goemans (MIT), Anupam Gupta (CMU), Nicole Immorlica (Northwestern), Michael Rabin (Harvard), and Eva Tardos (Cornell) -- kept it lively with questions and discussions.
Part of what makes Google ad systems exciting is the challenges they pose and overcome, and yet others in the horizon. It is remarkable how some of the fundamental problems Google ad systems grapple with are also some of the hardest research problems in the community of Market Algorithms. Joint Google and Academia meetings like this help researchers begin to attack these problems, and may be a model for research collaborations in the future.
Google University Research Awards
Posted on 09:56 by Unknown
Posted by Juan Vargas, University Relations
For most people, the word "Google" evokes associations of Internet search, free apps, and advertising--a far cry from our roots in research and academia. Google, however, has not lost sight of that fundamental relationship. Many of our offices (or "campuses") are located near local universities, and we maintain a working environment many consider to be collegiate: informal, collaborative, and home to expert lectures not only about science and technology, but also about literature, the economy, world peace, green energy, fitness, etc.
But these aspects are only one component of our enduring partnership with academia. Given the unique technical challenges we face, from the beginning we've often recognized the need for a strong relationship with the research community. One of the key ways we interact with this community is through the Google Research Awards Program.
Started in 2005, the program seeks to identify and support leading-edge research in strategic areas of engineering and computer science. Professors from universities worldwide submit proposals three times per year that are evaluated by a team of Google engineers and scientists.
Like Google itself, the program is global in scope: in the most recent round of submissions, nearly a third came from outside of the United States. During that round, we received 149 proposals and we ultimately decided to fund about a third of the projects.
A challenging aspect of running an awards program is making the funding decisions. We want to ensure that every proposal is reviewed by the best experts in the field. Luckily, we're fortunate to have many people at Google with the right expertise who help us make those decisions.
As Alfred Spector, VP for Research, explains, "The winning proposals not only get financial support from Google, but also receive the extra benefit of being assigned a Google liaison who maintains a special relation with the professor during the life of the award, contributes to the research, and ensures that the outcomes are valuable to Google and academia." One example of this is the recent paper featured on the Google Research home page, by Phil Long and Rocco Servedio from Columbia University titled "Random Classification Noise Defeats All Convex Potential Boosters."
The following are some highlights from other recently-funded projects:
Social Networks Research Through CourseRank
Hector Garcia-Molina, Stanford University
Hector Garcia-Molina and his team at Stanford are pursuing research on social networks and web usability by using and expanding CourseRank, a course evaluation and recommendation system for the Stanford community. As of June 2008, CourseRank has over 6,700 users and over 134,000 course evaluations. In addition to providing a service for Stanford students, it provides useful data to learn about social networks. Dr. Garcia-Molina plans to use CourseRank to investigate problems such as spam, trust, recommendations with complex objects, and the nature of social interactions. You can see a video with student testimonials and a demo here.
Energy-Efficient Storage Architectures for Data Centers
Drs. Gurumurthi and Stan from the University of Virgina have developed a new disk drive architecture called "Intra-Disk Parallelism" that they hope will reduce the energy consumption of data center storage systems by over 60% while providing high performance to applications. This architecture extends conventional disk drives by allowing it to handle multiple I/O requests in parallel and by provisioning additional hardware resources to enhance the parallel capabilities even further. Details of their research were published at the 2008 International Symposium on Computer Architecture (ISCA). The conference paper has also been selected to appear in the IEEE Micro special issue on Top Picks from Computer Architecture Conferences.
Finding Better Spoken Dialog System Metrics
When the number of calls to spoken dialog systems mounts into the thousands, or millions, it is impossible to listen to every call. Prof. Eskenazi and her team at CMU are studying new metrics that could improve the performance of those systems. Dr. Eskenazi hopes this research will help illuminate how best to cope with advances in tuned speech recognition and new synthetic voices.
Interested in learning more about the Google University Relations and our Research Awards Program? Please visit our web site where scholars can learn more and submit proposals.
Monday, 19 January 2009
Smart Thumbnails on YouTube
Posted on 17:35 by Unknown
Posted by Tomáš Ižo (Software Engineer) and Jay Yagnik (Head of Computer Vision Research)
One of our favorite aspects of Google web search are the informative snippets that appear with each search result. The more relevant they are, the quicker we can find what we're looking for. As members of the computer vision team, we're pleased to say that we now apply a similar philosophy to choosing thumbnails for videos on YouTube. After all, the thumbnail is the first visual piece of information our users get when searching or browsing YouTube videos. As we recently announced in a post on the YouTube Blog, our previous system of choosing thumbnails from the 25, 50 and 75% marks in the video, which often led to arbitrary, uninformative or sometimes even misleading images, is now a thing of the past. When a new video comes to YouTube, we now analyze it with an algorithm whose aim is to pick a set of images that are visually representative of the content of the video. As with the ground-breaking video identification tools that we launched on YouTube last year, this system is another example of how looking inside the video can lead to a safer, more relevant experience for our users.
Launching a computer vision algorithm on YouTube comes with a unique set of challenges, many of which have to do with the incredible diversity of the content and the awe-inspiring rate at which it pours into our servers: each minute YouTube receives as much as 13 hours of video content in a wide variety of genres, styles, formats and resolutions. Of course, even with some of the technical challenges aside, our work is far from finished. We'll keep on studying how our users interact with thumbnails and, more generally, with video content on the web, and we'll continue thinking of innovative ways to use computer vision and machine learning to improve the experience.
Subscribe to:
Posts (Atom)