Compact System

  • Subscribe to our RSS feed.
  • Twitter
  • StumbleUpon
  • Reddit
  • Facebook
  • Digg

Thursday, 25 April 2013

Two Googlers elected to the American Academy of Arts and Sciences

Posted on 13:52 by Unknown
Posted by Alfred Spector, Vice President, Engineering

Cross-posted with the Official Google Blog

On Wednesday, the American Academy of Arts and Sciences announced its list of 2013 elected members. We’re proud to congratulate Peter Norvig, director of research, and Arun Majumdar, vice president for energy; two Googlers who are among the new members elected this year.

Membership in the American Academy of Arts and Sciences is considered one of the nation’s highest honors, with those elected recognized as leaders in the arts, public affairs, business, and academic disciplines. With more than 250 Nobel Prize laureates and 60 Pulitzer Prize winners among its fellows, the American Academy celebrates the exceptional contributions of the elected members to critical social and intellectual issues.

With their election, Peter and Arun join six other Googlers as American Academy members: Eric Schmidt, Vint Cerf, Alfred Spector, Hal Varian, Ray Kurzweil, and founders Sergey Brin and Larry Page, all of whom embody our commitment to innovation and real-world impact. You can read more detailed summaries of Peter and Arun’s achievements below.

Dr. Peter Norvig, currently director of research at Google, is known most for his broad expertise in computer science and artificial intelligence, exemplified by his co-authorship (with Stuart Russell) of the leading college text, Artificial Intelligence: A Modern Approach. With more than 50 publications and a plethora of webpages, essays and software programs on a wide variety of CS topics, Peter is a catalyst of fundamental research across a wide range of disciplines while remaining a hands-on scientist who writes his own code. Recently, he has taught courses on artificial intelligence and the design of computer programs via massively open online courses (MOOC). Learn more about Peter and his research on norvig.com.

Dr. Arun Majumdar leads Google.org’s energy initiatives and advises Google on its broader energy strategy. Prior to joining Google last year, he was the founding director of the U.S. Department of Energy's Advanced Research Projects Agency-Energy (ARPA-E), where he served from October 2009 until June 2012. Earlier, he was a professor of mechanical engineering as well as materials science and engineering at the University of California, Berkeley, and headed the Environmental Energy Technologies Division at the Lawrence Berkeley National Laboratory. He has published several hundred papers, patents, and conference proceedings. Find out more about Arun.
Read More
Posted in | No comments

Thursday, 11 April 2013

50,000 Lessons on How to Read: a Relation Extraction Corpus

Posted on 09:00 by Unknown
Posted by Dave Orr, Product Manager, Google Research

One of the most difficult tasks in NLP is called relation extraction. It’s an example of information extraction, one of the goals of natural language understanding. A relation is a semantic connection between (at least) two entities. For instance, you could say that Jim Henson was in a spouse relation with Jane Henson (and in a creator relation with many beloved characters and shows).

The goal of relation extraction is to learn relations from unstructured natural language text. The relations can be used to answer questions (“Who created Kermit?”), learn which proteins interact in the biomedical literature, or to build a database of hundreds of millions of entities and billions of relations to try and help people explore the world’s information.

To help researchers investigate relation extraction, we’re releasing a human-judged dataset of two relations about public figures on Wikipedia: nearly 10,000 examples of “place of birth”, and over 40,000 examples of “attended or graduated from an institution”. Each of these was judged by at least 5 raters, and can be used to train or evaluate relation extraction systems. We also plan to release more relations of new types in the coming months.

Each relation is in the form of a triple: the relation in question, called a predicate; the subject of the relation; and the object of the relation. In the relation “Stephen Hawking graduated from Oxford,” Stephen Hawking is the subject, graduated from is the relation, and Oxford University is the object. Subjects and objects are represented by their Freebase MID’s, and the relation is defined as a Freebase property. So in this case, the triple would be represented as:

"pred":"/education/education/institution"
"sub":"/m/01tdnyh"
"obj":"/m/07tgn"

Just having the triples is interesting enough if you want a database of entities and relations, but doesn’t make much progress towards training or evaluation a relation extraction system. So we’ve also included the evidence for the relation, in the form of a URL and an excerpt from the web page that our raters judged. We’re also including examples where the evidence does not support the relation, so you have negative examples for use in training better extraction systems. Finally, we included ID’s and actual judgments of individual raters, so that you can filter triples by agreement.

Gory Details

The corpus itself, extracted from Wikipedia, can be found here: https://code.google.com/p/relation-extraction-corpus/

The files are in JSON format. Each line is a triple with the following fields:

  • pred: predicate of a triple
  • sub: subject of a triple
  • obj: object of a triple
  • evidences: an array of evidences for this triple
    • url: the web page from which this evidence was obtained
    • snippet: short piece of text supporting the triple
  • judgments: an array of judgements from human annotators
    • rator: hash code of the identity of the annotator
    • judgment: judgement of the annotator. It can take the values "yes" or "no"

Here’s an example:


{"pred":"/people/person/place_of_birth","sub":"/m/026_tl9","obj":"/m/02_286","evidences":[{"url":"http://en.wikipedia.org/wiki/Morris_S._Miller","snippet":"Morris Smith Miller (July 31, 1779 -- November 16, 1824) was a United States Representative from New York. Born in New York City, he graduated from Union College in Schenectady in 1798. He studied law and was admitted to the bar. Miller served as private secretary to Governor Jay, and subsequently, in 1806, commenced the practice of his profession in Utica. He was president of the village of Utica in 1808 and judge of the court of common pleas of Oneida County from 1810 until his death."}],"judgments":[{"rater":"11595942516201422884","judgment":"yes"},{"rater":"16169597761094238409","judgment":"yes"},{"rater":"1014448455121957356","judgment":"yes"},{"rater":"16651790297630307764","judgment":"yes"},{"rater":"1855142007844680025","judgment":"yes"}]}

The web is chock full of information, put there to be read and learned from. Our hope is that this corpus is a small step towards computational understanding of the wealth of relations to be found everywhere you look.

This dataset is licensed by Google Inc. under the Creative Commons Attribution-Sharealike 3.0 license.

Thanks to Shaohua Sun, Ni Lao, and Rahul Gupta for putting this dataset together.

Thanks also to Michael Ringgaard, Fernando Pereira, Amar Subramanya, Evgeniy Gabrilovich, and John Giannandrea for making this data release possible.
Read More
Posted in Natural Language Processing, Wiki | No comments

Tuesday, 9 April 2013

Advanced Power Searching with Google: Lessons Learned

Posted on 09:30 by Unknown
Posted by Dan Russell, Uber Tech Lead, Search Quality & User Happiness and Maggie Johnson, Director of Education and University Relations

Large classes are something you normally want to avoid like the plague. So the idea of being in a class with tens of thousands of students seems like a completely crazy idea.

But in January, 2013, Google offered a free “MOOC” (a Massive Open Online Course) to teach Advanced Power Searching (APS) to a wide variety of information professionals.

The wholly online class ran for two weeks covering advanced research skills in a challenge-based format. It also had a bit more than 35,000 students sign up for the class.

In this case, the large class size was a boon to the students. Not only was there a vigorous discussion of the material in the social media, but with a class this large, anytime you had a question, someone else in the class had almost certainly asked the same question and had an answer ready. As in many MOOCs, the large online class size did not stress any lecture hall capacities, but it did give the students the benefit of multicultural classmates that were effectively always present in the social spaces of the MOOC.

A typical Massive Open Online Course (MOOC) is a simple progression through a series of mini-lectures--usually a short video followed by reflective questions, problem sets and a few assessments. MOOCs can have huge numbers of students; dozens have been offered with over 150,000 students enrolled. Based on our experiments with Power Searching with Google in 2012, we wanted to do something different. When we offered Advanced Power Searching with Google (APS) in January of 2013, we decided to try out a number of new ideas.

Through this course, we wanted to enable our students to solve complex research questions using a variety of tools, such as Google Scholar, Patents, Books, Google+, etc.. We defined complex problems that had more than one right answer and more than one way to find those answers.

Unlike a traditional MOOC, the APS course had twelve challenges that students could tackle in any order they liked. There were four easy, four medium and four difficult challenges. Part of the design of the class was to have students discover the skills they’d need to solve the challenges and select appropriate video or text lessons. Students could also access case studies that showed how others solve similar problems.

We called our MOOC design “Choose your own adventure.” Each challenge presented a research question like this:


“You are in the city that is home to the House of Light. Nearby there is a museum in a converted school featuring paintings from the far-away Forest of Honey.


What traditional festival are you visiting?”


In this class, the large cohort of 35,000 students worked through the materials together, using online forums to ask questions as well as Google+ Hangouts to attend office hours and collaborate on solving challenges. Instructor Dan Russell and a group of teaching assistants monitored students’ activities and provided support as needed.

If they needed additional help, students could post a question on the forum or see how others solved the challenge. Students could post their solutions to challenges in a special “Peer explanations” section; a feature that many students appreciated as it let them see how others in the class approached the problem in their own ways.

In analyzing the data, we found that there were a decreasing number of views on each challenge page, indicating that students most likely tried the challenges in the order given. While some liked the ability to jump around, most tended to go through the content linearly. Most students who completed the course tried (or at least looked at) all twelve challenges. Many students who did not complete the course tried three or fewer challenges.

To earn a certificate of completion, students submitted two detailed case studies of how they solved a complex search challenge. Students provided great examples of how they used Google tools to research their family’s history, the origins of common objects, or trips they anticipate taking. In addition to listing their queries, they wrote details about how they knew websites were credible and what they learned along the way.

To assess their work, we experimented with letting the students grade their assignments based on a rubric. We collected their scores and compared them with a random sample of assignments graded by TAs. There was a moderate yet statistically significant correlation (r=0.44) between student scores and TA scores. In fact, the majority of students graded themselves within two points of how an expert grader assessed their work. This is a positive result since it suggests that self-graded project work in a MOOC can be valuable as a source of insight into student performance.

The challenge format seemed to be effective and motivating for a small, dedicated population of students. We had 35,000 registrants for this advanced course, and 12% earned a certificate of completion. This rate is somewhat lower than what we saw for Power Searching with Google, a more traditional MOOC. Students who did not complete the course reported a lack of time, and difficulty of the content as barriers.

One interesting point was that labeling the challenges as easy, medium or difficult likely had an unintentional effect. The first challenge was marked as “easy,” but many people found it difficult. This may have de-motivated students from attempting more difficult challenges. Next time, we plan to ask students if the first challenge was too easy, or too challenging, and then send them to a challenge at an appropriate level of difficulty.

Watch for more MOOCs on our products and services in the coming months. And watch for more experimentation as we apply what we have learned, and try more ideas and new approaches in future online courses.
Read More
Posted in Education, MOOC | No comments

Wednesday, 27 March 2013

Education Awards on Google App Engine

Posted on 10:00 by Unknown
Posted by Andrea Held, Google University Relations

Cross-posted with Google Developers Blog

Last year we invited proposals for innovative projects built on Google’s infrastructure. Today we are pleased to announce the 11 recipients of a Google App Engine Education Award. Professors and their students are using the award in cloud computing courses to study databases, distributed systems, web mashups and to build educational applications. Each selected project received $1000 in Google App Engine credits.

Awarding computational resources to classroom projects is always gratifying. It is impressive to see the creative ideas students and educators bring to these programs.
Below is a brief introduction to each project. Congratulations to the recipients!

John David N. Dionisio, Loyola Marymount University
Project description: The objective of this undergraduate database systems course is for students to implement one database application in two technology stacks, a traditional relational database and on Google App Engine. Students are asked to study both models and provide concrete comparison points.

Xiaohui (Helen) Gu, North Carolina State University
Project description: Advanced Distributed Systems Class
The goal of the project is to allow the students to learn distributed system concepts by developing real distributed system management systems and testing them on real world cloud computing infrastructures such as Google App Engine.

Shriram Krishnamurthi, Brown University
Project description: WeScheme is a programming environment that runs in the Web browser and supports interactive development. WeScheme uses App Engine to handle user accounts, serverside compilation, and file management.

Feifei Li, University of Utah
Project description: A graduate-level course that will be offered in Fall 2013 on the design and implementation of large data management system kernels. The objective is to integrate features from a relational database engine with some of the new features from NoSQL systems to enable efficient and scalable data management over a cluster of commodity machines.

Mark Liffiton, Illinois Wesleyan University
Project description: TeacherTap is a free, simple classroom-response system built on Google App Engine. It lets students give instant, anonymous feedback to teachers about a lecture or discussion from any computer or mobile device with a web browser, facilitating more adaptive class sessions.

Eni Mustafaraj, Wellesley College
Project description: Topics in Computer Science: Web Mashups. A CS2 course that combines Google App Engine and MIT App Inventor. Students will learn to build apps with App Inventor to collect data about their life on campus. They will use Google App Engine to build web services and apps to host the data and remix it to create web mashups. Offered in the 2013 Spring semester.

Manish Parashar, Rutgers University
Project description: Cloud Computing for Scientific Applications -- Autonomic Cloud Computing teaches students how a hybrid HPC/Grid + Cloud cyber infrastructure can be effectively used to support real-world science and engineering applications. The goal of our efforts is to explore application formulations, Cloud and hybrid HPC/Grid + Cloud infrastructure usage modes that are meaningful for various classes of science and engineering application workflows.

Orit Shaer, Wellesley College
Project description: GreenTouch
GreenTouch is a collaborative environment that enables novice users to engage in authentic scientific inquiry. It consists of a mobile user interface for capturing data in the field, a web application for data curation in the cloud, and a tabletop user interface for exploratory analysis of heterogeneous data.

Elliot Soloway, University of Michigan
Project description: WeLearn Mobile Platform: Making Mobile Devices Effective Tools for K-12. The platform makes mobile devices (Android, iOS, WP8) effective, essential tools for all-the-time, everywhere learning. WeLearn’s suite of productivity and communication apps enable learners to work collaboratively; WeLearn’s portal, hosted on Google App Engine, enables teachers to send assignments, review, and grade student artifacts. WeLearn is available to educators at no charge.

Jonathan White, Harding University
Project description: Teaching Cloud Computing in an Introduction to Engineering class for freshmen. We explore how well-designed systems are built to withstand unpredictable stresses, whether that system is a building, a piece of software or even the human body. The grant from Google is allowing us to add an overview of cloud computing as a platform that is robust under diverse loads.

Dr. Jiaofei Zhong, University of Central Missouri
Project description: By building an online Course Management System, students will be able to work on their team projects in the cloud. The system allows instructors and students to manage the course materials, including course syllabus, slides, assignments and tests in the cloud; the tool can be shared with educational institutions worldwide.

Read More
Posted in App Engine, University Relations | No comments

Wednesday, 13 March 2013

Scaling Computer Science Education

Posted on 11:15 by Unknown
Posted by Maggie Johnson, Director of Education and University Relations

Last week, I attended the annual SIGCSE (Special Interest Group, Computer Science Education) conference in Denver, CO. Google has been a platinum sponsor of SIGCSE for many years now, and the conference provides an opportunity for hundreds of computer science (CS) educators to share ideas and work on strategies to bring high quality CS education to K12 and undergraduate students.

Significant accomplishments over the last few years have laid a strong foundation for scaling CS curriculum, professional development (PD) and related programs in this country. The NSF has been funding curriculum and PD around the new CS Principles Advanced Placement course. The CSTA has published standards for K12 CS and a report on the limited extent to which schools, districts and states provide CS instruction to their students. CS Advocacy group, Computing in the Core, even provides a toolkit for communities to follow as they urge legislators for integration of Computer Science education into core K12 curriculum.

All of this work has made an impact, but there is still more to do.

I see our priorities in CS education to be ones of awareness and access. As CS educators, we must continue to raise awareness about the tremendous demand for jobs in the computing sector, and balance misconceptions with accurate data. Many students, parents, teachers and administrators remember the hype and disillusionment of the Dotcom period and myths on outsourcing and dwindling jobs yet the US Bureau of Labor Statistics (BLS) reports that ⅔ of all job growth in Science and Engineering will be in Computer Science employment over the next decade. (See 2010 BLS report here.) Clearing up this misconception is essential if we hope to satisfy US labor needs with recent graduates over the next several years.
Source: Gianchandani, Erwin. Revisiting ‘Where the Jobs Are’. The Computing Community Consortium Blog post on 23 May 2012. Link accessed on 8 March 2013.

Another misconception surrounds the range of CS-focused occupations that exist. The world of CS is expanding rapidly and we should celebrate the diversity of CS applications that are gaining momentum. Instead of the archetype of a sun-starved computer scientist, or software engineers working in isolation with little teamwork or communication opportunities, educators can encourage project-based learning, video game development, robotics, and graphic design as more concrete representations for abstract computational thinking.

Google believes that computing and CS are critical to our future, not only in the high tech sector, but for everyone. Our economy is becoming more and more dependent on technology-based solutions, which will require a future workforce with significant levels of CS knowledge and experience. In addition, we anticipate new career opportunities opening up in the next 3-5 years as more businesses move into the cloud and shift the way they run their IT departments.

Help us get the word out about the great opportunities in computing through organizations such as code.org, ACM, and NCWIT. Google is doing its part to support CS education and outreach through many programs including CS4HS, our Exploring Computational Thinking curriculum, and several student and teacher programs. So much opportunity, so little time!
Read More
Posted in Computer Science | No comments

Tuesday, 12 March 2013

Our Commitment to Social Computing Research: Social Interactions Focused Awards Announcement

Posted on 09:00 by Unknown
Ed H. Chi, Staff Research Scientist

Social interactions have always been an important part of the human experience. Social interaction research has shown results ranging from influences on our behavior from social networks [Aral2012] to our understanding of social belonging on health [Walton2011], as well as how conflicts and coordination play out in Wikipedia [Kittur2007]. Interestingly, social scientists have studied social interactions for many years, but it wasn’t until very recently that researchers can study these mechanisms through the explosion of services and data available on web-based social systems.

From information dissemination and the spread of innovation and ideas, to scientific discovery, we are seeing how a deep understanding of social interactions is affecting many different fields, such as health and education. For instance, scientists now have strong evidence that social interactions underlie many fundamental learning mechanisms starting from infancy well into adulthood [Meltzoff2009], and that peer discussions are critical in conceptual learning in college classes [Smith2009]. How might these learning science findings be built into social systems and products so that users maximize what they learn on the Web?

We know that interactions on the Web are diverse and people-centered. Google now enables social interactions to occur across many of our products, from Google+ to Search to YouTube. To understand the future of this socially connected web, we need to investigate fundamental patterns, design principles, and laws that shape and govern these social interactions.

We envision research at the intersection of disciplines including Computer Science, Human-Computer Interaction (HCI), Social Science, Social Psychology, Machine Learning, Big Data Analytics, Statistics and Economics. These fields are central to the study of how social interactions work, particularly driven by new sources of data, for example, open data sets from Web2.0 and social media sites, government databases, crowdsourcing, new survey techniques, and crisis management data collections. New techniques from network science and computational modeling, social network and sentiment analysis, application of statistical and machine learning, as well as theories from evolutionary theory, physics, and information theory, are actively being used in social interaction research.

We’re pleased to announce that Google has awarded over $1.2 million dollars to support the Social Interactions Research Awards, which are given to university research groups doing work in social computing and interactions. Research topics range from crowdsourcing, social annotations, a social media behavioral study, social learning, conversation curation, and scientific studies of how to start online communities.

We have awarded 15 researchers in 7 universities. We selected these proposals after a rigorous internal review. We believe the results will be broadly useful to product development and will further scientific research.

  • Joseph Konstan, Loren Terveen, and John Riedl from University of Minnesota. Precision Crowdsourcing: Closing the Loop to turn Information Consumers into Information Contributors.
  • Mor Naaman from Rutgers University, and Oded Nov from Polytechnic Institute of New York University. Examining the Impact of Social Traces on Page Visitors’ Opinions and Engagement.
  • Paul Resnick, Eytan Adar, and Cliff Lampe from University of Michigan. MTogether: A Living Lab for Social Media Research.
  • Marti Hearst from UC Berkeley. Understanding Social Learning Among Subgroups Within Large Online Learning Environments.
  • David Karger and Rob Miller from MIT. Crowdsourced Curation of Conversations.
  • Robert Kraut, Laura Dabbish, Jason Hong, Aniket Kittur from CMU. Successfully Starting Online Groups.

We look forward to working with these researchers, and we hope that we will jointly push the frontier of social interactions research to the next level.

References
[1] Aral, S., & Walker, D. (2012). Identifying Influential and Susceptible Members of Social Networks. Science , 337 (6092 ), 337–341. doi:10.1126/science.1215842
[2] Walton, G. M., & Cohen, G. L. (2011). A Brief Social-Belonging Intervention Improves Academic and Health Outcomes of Minority Students. Science , 331 (6023 ), 1447–1451. doi:10.1126/science.1198364
[3] Aniket Kittur, Bongwon Suh, Bryan Pendleton, Ed H. Chi. He Says, She Says: Conflict and Coordination in Wikipedia. In Proc. of ACM Conference on Human Factors in Computing Systems (CHI2007), pp. 453--462, April 2007. ACM Press. San Jose, CA.
[4] Meltzoff, A. N., Kuhl, P. K., Movellan, J., & Sejnowski, T. J. (2009). Foundations for a New Science of Learning. Science , 325 (5938), 284–288. doi:10.1126/science.1175626
[5] Smith, M. K., Wood, W. B., Adams, W. K., Wieman, C., Knight, J. K., Guild, N., & Su, T. T. (2009). Why Peer Discussion Improves Student Performance on In-Class Concept Questions. Science , 323 (5910), 122–124. doi:10.1126/science.1165919
Read More
Posted in Research Awards, University Relations | No comments

Friday, 8 March 2013

Learning from Big Data: 40 Million Entities in Context

Posted on 10:30 by Unknown
Posted by Dave Orr, Amar Subramanya, and Fernando Pereira, Google Research

When someone mentions Mercury, are they talking about the planet, the god, the car, the element, Freddie, or one of some 89 other possibilities? This problem is called disambiguation (a word that is itself ambiguous), and while it’s necessary for communication, and humans are amazingly good at it (when was the last time you confused a fruit with a giant tech company?), computers need help.

To provide that help, we are releasing the Wikilinks Corpus: 40 million total disambiguated mentions within over 10 million web pages -- over 100 times bigger than the next largest corpus (about 100,000 documents, see the table below for mention and entity counts). The mentions are found by looking for links to Wikipedia pages where the anchor text of the link closely matches the title of the target Wikipedia page. If we think of each page on Wikipedia as an entity (an idea we’ve discussed before), then the anchor text can be thought of as a mention of the corresponding entity.

Dataset Number of Mentions Number of Entities
Bentivogli et al. (data) (2008) 43,704 709
Day et al. (2008) less than 55,0003,660
Artiles et al. (data) (2010) 57,357 300
Wikilinks Corpus 40,323,863 2,933,659

What might you do with this data? Well, we’ve already written one ACL paper on cross-document co-reference (and received lots of requests for the underlying data, which partly motivates this release). And really, we look forward to seeing what you are going to do with it! But here are a few ideas:
  • Look into coreference -- when different mentions mention the same entity -- or entity resolution -- matching a mention to the underlying entity
  • Work on the bigger problem of cross-document coreference, which is how to find out if different web pages are talking about the same person or other entity
  • Learn things about entities by aggregating information across all the documents they’re mentioned in
  • Type tagging tries to assign types (they could be broad, like person, location, or specific, like amusement park ride) to entities. To the extent that the Wikipedia pages contain the type information you’re interested in, it would be easy to construct a training set that annotates the Wikilinks entities with types from Wikipedia.
  • Work on any of the above, or more, on subsets of the data. With existing datasets, it wasn’t possible to work on just musicians or chefs or train stations, because the sample sizes would be too small. But with 10 million Web pages, you can find a decent sampling of almost anything.

Gory Details

How do you actually get the data? It’s right here: Google’s Wikilinks Corpus. Tools and data with extra context can be found on our partners’ page: UMass Wiki-links. Understanding the corpus, however, is a little bit involved.

For copyright reasons, we cannot distribute actual annotated web pages. Instead, we’re providing an index of URLs, and the tools to create the dataset, or whichever slice of it you care about, yourself. Specifically, we’re providing:
  • The URLs of all the pages that contain labeled mentions, which are links to English Wikipedia
  • The anchor text of the link (the mention string), the Wikipedia link target, and the byte offset of the link for every page in the set
  • The byte offset of the 10 least frequent words on the page, to act as a signature to ensure that the underlying text hasn’t changed -- think of this as a version, or fingerprint, of the page
  • Software tools (on the UMass site) to: download the web pages; extract the mentions, with ways to recover if the byte offsets don’t match; select the text around the mentions as local context; and compute evaluation metrics over predicted entities.
The format looks like this:

URL http://1967mercurycougar.blogspot.com/2009_10_01_archive.html
MENTION Lincoln Continental Mark IV 40110 http://en.wikipedia.org/wiki/Lincoln_Continental_Mark_IV
MENTION 1975 MGB roadster 41481 http://en.wikipedia.org/wiki/MG_MGB
MENTION Buick Riviera 43316 http://en.wikipedia.org/wiki/Buick_Riviera
MENTION Oldsmobile Toronado 43397 http://en.wikipedia.org/wiki/Oldsmobile_Toronado
TOKEN seen 58190
TOKEN crush 63118
TOKEN owners 69290
TOKEN desk 59772
TOKEN relocate 70683
TOKEN promote 35016
TOKEN between 70846
TOKEN re 52821
TOKEN getting 68968
TOKEN felt 41508


We’d love to hear what you’re working on, and look forward to what you can do with 40 million mentions across over 10 million web pages!

Thanks to our collaborators at UMass Amherst: Sameer Singh and Andrew McCallum.

Read More
Posted in Natural Language Processing, wikipedia | No comments
Newer Posts Older Posts Home
Subscribe to: Posts (Atom)

Popular Posts

  • Our Faculty Institute brings faculty back to the drawing board
    Posted by Nina Kim Schultz, Google Education Research Cross-posted with the Official Google Blog School may still be out for summer, but tea...
  • Academic Successes in Cluster Computing
    Posted by Alfred Spector, VP of Research Access to massive computing resources is foundational to Research and Development. Fifteen awardees...
  • Towards Energy-Proportional Datacenters
    Posted by Dennis Abts, Michael R. Marty, Philip M. Wells, Peter Klausler, and Hong Liu This is part of the series highlighting some notable...
  • International Conference on Machine Learning (ICML 2009) in Montreal
    Posted by Eyal Even Dar and Vahab Mirrokni , Google Research, NY The 26th International Conference on Machine Learning ( ICML 2009 ) was re...
  • A new landmark in computer vision
    Posted by Jay Yagnik, Head of Computer Vision Research [Cross-posted with the Official Google Blog ] Science fiction books and movies have l...
  • Market Algorithms and Optimization Meeting
    Posted by  Vahab S. Mirrokni and Muthu Muthukrishnan Google auctions ads, and enables a market with millions of advertisers and users.  This...
  • Education Awards on Google App Engine
    Posted by Andrea Held, Google University Relations Cross-posted with Google Developers Blog Last year we invited proposals for innovative p...
  • Speed Matters
    Posted by Jake Brutlag, Web Search Infrastructure At Google, we've gathered hard data to reinforce our intuition that "speed matter...
  • Google launches Korean Voice Search
    Posted by Mike Schuster & Martin Jansche, Google Research On June 16th, we launched our Korean voice search system . Google Search by Vo...
  • Two Views from the 2009 Google Faculty Summit
    Posted by Alfred Spector, Vice President of Research and Special Initiatives [cross-posted with the Official Google Blog ] We held our fifth...

Categories

  • accessibility
  • ACL
  • ACM
  • Acoustic Modeling
  • ads
  • adsense
  • adwords
  • Africa
  • Android
  • API
  • App Engine
  • App Inventor
  • Audio
  • Awards
  • Cantonese
  • China
  • Computer Science
  • conference
  • conferences
  • correlate
  • crowd-sourcing
  • CVPR
  • datasets
  • Deep Learning
  • distributed systems
  • Earth Engine
  • economics
  • Education
  • Electronic Commerce and Algorithms
  • EMEA
  • EMNLP
  • entities
  • Exacycle
  • Faculty Institute
  • Faculty Summit
  • Fusion Tables
  • gamification
  • Google Books
  • Google+
  • Government
  • grants
  • HCI
  • Image Annotation
  • Information Retrieval
  • internationalization
  • Interspeech
  • jsm
  • jsm2011
  • K-12
  • Korean
  • Labs
  • localization
  • Machine Hearing
  • Machine Learning
  • Machine Translation
  • MapReduce
  • market algorithms
  • Market Research
  • ML
  • MOOC
  • NAACL
  • Natural Language Processing
  • Networks
  • Ngram
  • NIPS
  • NLP
  • open source
  • operating systems
  • osdi
  • osdi10
  • patents
  • ph.d. fellowship
  • PiLab
  • Policy
  • Public Data Explorer
  • publication
  • Publications
  • renewable energy
  • Research Awards
  • resource optimization
  • Search
  • search ads
  • Security and Privacy
  • SIGMOD
  • Site Reliability Engineering
  • Speech
  • statistics
  • Structured Data
  • Systems
  • Translate
  • trends
  • TV
  • UI
  • University Relations
  • UNIX
  • User Experience
  • video
  • Vision Research
  • Visiting Faculty
  • Visualization
  • Voice Search
  • Wiki
  • wikipedia
  • WWW
  • YouTube

Blog Archive

  • ▼  2013 (51)
    • ▼  December (3)
      • Groundbreaking simulations by Google Exacycle Visi...
      • Googler Moti Yung elected as 2013 ACM Fellow
      • Free Language Lessons for Computers
    • ►  November (9)
    • ►  October (2)
    • ►  September (5)
    • ►  August (2)
    • ►  July (6)
    • ►  June (7)
    • ►  May (5)
    • ►  April (3)
    • ►  March (4)
    • ►  February (4)
    • ►  January (1)
  • ►  2012 (59)
    • ►  December (4)
    • ►  October (4)
    • ►  September (3)
    • ►  August (9)
    • ►  July (9)
    • ►  June (7)
    • ►  May (7)
    • ►  April (2)
    • ►  March (7)
    • ►  February (3)
    • ►  January (4)
  • ►  2011 (51)
    • ►  December (5)
    • ►  November (2)
    • ►  September (3)
    • ►  August (4)
    • ►  July (9)
    • ►  June (6)
    • ►  May (4)
    • ►  April (4)
    • ►  March (5)
    • ►  February (5)
    • ►  January (4)
  • ►  2010 (44)
    • ►  December (7)
    • ►  November (2)
    • ►  October (9)
    • ►  September (7)
    • ►  August (2)
    • ►  July (7)
    • ►  June (3)
    • ►  May (2)
    • ►  April (1)
    • ►  March (1)
    • ►  February (1)
    • ►  January (2)
  • ►  2009 (44)
    • ►  December (8)
    • ►  November (4)
    • ►  August (4)
    • ►  July (5)
    • ►  June (5)
    • ►  May (4)
    • ►  April (6)
    • ►  March (3)
    • ►  February (1)
    • ►  January (4)
  • ►  2008 (11)
    • ►  December (1)
    • ►  November (1)
    • ►  October (1)
    • ►  September (1)
    • ►  July (1)
    • ►  May (3)
    • ►  April (1)
    • ►  March (1)
    • ►  February (1)
  • ►  2007 (9)
    • ►  October (1)
    • ►  September (2)
    • ►  August (1)
    • ►  July (1)
    • ►  June (2)
    • ►  February (2)
  • ►  2006 (15)
    • ►  December (1)
    • ►  November (1)
    • ►  September (1)
    • ►  August (1)
    • ►  July (1)
    • ►  June (2)
    • ►  April (3)
    • ►  March (4)
    • ►  February (1)
Powered by Blogger.

About Me

Unknown
View my complete profile