Compact System

  • Subscribe to our RSS feed.
  • Twitter
  • StumbleUpon
  • Reddit
  • Facebook
  • Digg

Monday, 19 January 2009

Smart Thumbnails on YouTube

Posted on 17:35 by Unknown
Posted by Tomáš Ižo (Software Engineer) and Jay Yagnik (Head of Computer Vision Research)

One of our favorite aspects of Google web search are the informative snippets that appear with each search result.  The more relevant they are, the quicker we can find what we're looking for.  As members of the computer vision team, we're pleased to say that we now apply a similar philosophy to choosing thumbnails for videos on YouTube.  After all, the thumbnail is the first visual piece of information our users get when searching or browsing YouTube videos.  As we recently announced in a post on the YouTube Blog, our previous system of choosing thumbnails from the 25, 50 and 75% marks in the video, which often led to arbitrary, uninformative or sometimes even misleading images, is now a thing of the past.  When a new video comes to YouTube, we now analyze it with an algorithm whose aim is to pick a set of images that are visually representative of the content of the video.  As with the ground-breaking video identification tools that we launched on YouTube last year, this system is another example of how looking inside the video can lead to a safer, more relevant experience for our users.  

Launching a computer vision algorithm on YouTube comes with a unique set of challenges, many of which have to do with the incredible diversity of the content and the awe-inspiring rate at which it pours into our servers: each minute YouTube receives as much as 13 hours of video content in a wide variety of genres, styles, formats and resolutions.  Of course, even with some of the technical challenges aside, our work is far from finished.  We'll keep on studying how our users interact with thumbnails and, more generally, with video content on the web, and we'll continue thinking of innovative ways to use computer vision and machine learning to improve the experience.
Read More
Posted in | No comments

Monday, 12 January 2009

Maybe your computer just needs a hug

Posted on 07:45 by Unknown
Posted by Anthony Francis

Maybe your computer just needs a hug
I'm Anthony Francis, an artificial intelligence researcher working in Google's Search Quality group.  One of the things I like about Google is that we give back to the world in ways that make sense for our business, both as a company and as individuals.  Google's search engine runs on electrical power, and so Google as a company invests in renewable energy and encourages Googlers as individuals to conserve energy.  Google's search engine runs on open source software, and so Google as a company open sources its software and encourages Googlers to participate in existing open source projects.

What may be less obvious is that Google's search engine runs on ideas.  Google as a company has a great Research department, but it also encourages Googlers to participate in the research community - both in visible, obvious ways like writing papers and attending conferences, but also in less visible but still necessary chores that keep the research community running, like 
reviewing papers and organizing conferences.  What I'd like to tell you about today is how Google helped me give back to the research community by letting me take time to write a paper on work I did before I came to Google.

The Personal Pet Project
Ashwin Ram's Cognitive Computing Lab at Georgia Tech has been working on emotional agents for over ten years, first with robots, and more recently in computer games and virtual environments.  I worked on one of our earliest efforts in this area, the PEPE (Personal Pet) project, a joint effort with Yamaha to develop an robotic pet.  Dr. Ram noticed that many consumer electronics require configuration steps that are frustrating, even baffling: everything from VCRs to toasters has blinking clocks and unused features.  Pets, on the other hand, don't have to be "configured": they understand us on emotional level, noticing what makes us happy or angry, remembering what brings reward and punishment, and learning to respond appropriately without us reading a manual or pushing a button.

PEPE emulated this intuitive understanding with an emotional long term memory module, which I developed.  This module had three parts: basic emotions, emotional memory, and emotional reminding.  Basic emotions gave the robot realistic behavior "out of the box": for example, the robot interpreted being petted on the head as being pleasant, which made it want to socialize; it interpreted a kick to the rear as unpleasant, which made it want to hide.  The emotional memory associated these primitive emotions with the people and objects the robot saw in its environment.  Emotional reminding closed the loop: when the robot saw a person or object again, it was reminded of the past emotion, which influenced but did not determine the robot's ultimate emotional expression.  So the robot naturally wanted to play with people who petted it and fled people who had kicked it in the past, but it would be possible to overcome that past learning by ignoring the robot (if you didn't want to play with it as much anymore) or by petting it (if you had kicked it but no longer wanted it to be scared of you).

I worked with a team of engineers at Yamaha in Japan to implement this emotional long term memory on a small robot they were building.  Our job was easier because PEPE's basic architecture enabled creating many different kinds of behaviors that could easily interact with each other, in parallel or in sequence.  On top of this architecture, we added basic emotions, emotional learning and emotional reminding in stages.  Our initial tests were positive, and upon my return to the United States, we began porting this software back to Georgia Tech's PEPE robot.  PEPE even had a brief moment of fame, appearing on a local TV station; ultimately, however, Georgia Tech and Yamaha decided to move on to new projects.

Characters with Personality
But the dream of emotional agents didn't die.  Manish Mehta, a graduate student at Georgia Tech researching interactive games, began working with Ram to try to develop more sophisticated models of emotional change.  Mehta realized that emotional events can act as a trigger for behavioral change: if we humans try something that works really poorly, or spectacularly well, the emotions generated by that experience can prompt us to make changes to our behavior in the future --- for example, remembering to bring that umbrella in the future so we no longer have to walk home wet in the rain.

Mehta used ABL, an agent language developed by Michael Mateas and Andrew Stern for the Facade project.  Like the fundamental architecture of PEPE, ABL enabled the author of an agent to design complicated behaviors that could interact in sequence, in parallel, or through complex triggering.  Mehta extended ABL to allow an agent to rewrite its own behaviors, and triggered that rewriting based on charged emotional events.  This worked: using two agents playing tag, Mehta was able to make them revise their behaviors based on how well they played the game. However, it had an unintended consequence: one of the agents decided it was less stressed out when it was "IT".  When it was "IT", it wasn't being chased and wasn't feeling stressed, so it just let itself get tagged, enabling it to "chase" the other player at a leisurely pace.

While this result was surprising, even amazing, it failed to achieved the original objective.   Mehta realized this had exposed one more property of human behavioral change: emotional events can prompt change, but we also have a self image - a model of what we think our personality should be like.  As our personalities change, we audit the changes to make sure they fit our self image and apply further corrections.  For a computer game character, that "personality" is really the designer's intent, so Mehta augmented his emotional learning system to check potential behavioral changes and make sure that they did not violate the original intended design.  With this change, both characters were able to learn from their good and bad experiences playing tag, and neither "quit playing the game" just because they didn't want the stress of being chased.

Putting It All Together
The paper that Ram, and Mehta and I wrote for the Handbook of Synthetic Emotions and Sociable Robotics reports this work and ties together all of these threads.  We were not alone in challenges we faced in developing intelligent agents: others faced them too, in robotics, computer games and academic AI.  Many people use flexible behavior systems like PEPE's architecture or like ABL, but building large systems out of these can be challenging and extremely labor intensive.  What Mehta and I found, however, is that the use of emotion models acted like a force multiplier: adding basic emotional responses to an existing behavior system radically increased the flexibility and apparent realism of an intelligent agent, whether it was a robotic puppy or a character in a computer game.

Moreover, the more components of the emotion stack you add to an agent, the easier it becomes to extend.  While adding emotion made PEPE extended the behaviors it could perform; modifying those emotions using memories changed the character of its behavior, making it seem more lifelike.  Mehta's work in extending ABL took this further, enabling creative behavior changes in response to agent frustration.  Adding personality models makes these changes stable, enabling the agent to adapt to new situations while staying consistent with their designer's intent.  Often, just adding more behaviors to an existing agent can make it more brittle; we found in contrast adding new layers inspired by research in human and animal emotion made the agent more sophisticated and robust.

Giving Something Back
Writing this paper did not take much time from Google.  Most of it was done in the evening on my own time, with the occasional email exchange with my coauthors and editors during the day while I was waiting on compiles.  And that is how it should be: it may be a long time before this work directly benefits Google, since we do not currently develop robots or computer games.  But Google still benefited because the advances that we need to improve our search engine often start in academia.

By encouraging our staff to follow up on their research work and to contribute to the research community, Google supports the growth of the next big idea.  Knowing that we can follow up on our past work makes researchers at Google feel more empowered to pursue new ideas and continually exposes us to sources of inspiration.  And that inspiration, which makes Google a better place to work, is exactly what I need to keep trying new ideas, which in the end will help make Google a better place to search.
Read More
Posted in | No comments

Tuesday, 30 December 2008

Translation is Risky Business

Posted on 13:52 by Unknown
Posted by Shankar Kumar and Wolfgang Macherey

At Google, we like search. So it's no surprise that we treat language translation as a search problem. We build statistical models of how one language maps to another (the translation model) and models of what the target language is supposed to look like (the language model) and then we search for the best translation according to those models (combined into one big log linear model for those of you taking notes).

But, just as putting all of your money in the investment with the highest historical return is not always the best idea, choosing the translation with the highest probability is not always the best idea either - especially when you have a relatively flat distribution among the top candidates. Instead, we can use the Minimum Bayes Risk (MBR) criterion. Essentially, we look at a sample of the best candidate translations (the so called n-best list) and choose the safest one, the one most likely to do the least amount of damage (where 'damage' is defined by our measurement of translation quality). You might want to view this as choosing a translation that is a lot like the other good translations instead of choosing that strange one that had the good model score.

If this is our 'diversification' strategy, how can we make things even safer? Exactly the same way as we do for investments, we diversify even more. That is, we look at more of the candidate translations to make the MBR decision. A lot more. And the way to do that is to build a lattice of translations during the search and then we do our MBR search over the lattice. Instead of 100 or 1000 best translations that we would use for the n-best approach, lattices give us access to a number that rivals the number of particles in the visible universe (really, it's huge).

 Interested? You can read all about it here.
Read More
Posted in | No comments

Monday, 10 November 2008

plop: Probabilistic Learning of Programs

Posted on 16:11 by Unknown
Posted by Moshe Looks

Cross-posted with Open Source at Google blog

Traditional machine learning systems work with relatively flat, uniform data representations, such as feature vectors, time-series, and probabilistic context-free grammars. However, reality often presents us with data which are best understood in terms of relations, types, hierarchies, and complex functional forms. The best representational scheme we computer scientists have for coping with this sort of complexity is computer programs. Yet there are comparatively few machine learning methods that operate directly on programmatic representations, due to the extreme combinatorial explosions involved and the semantic complexities of programs.

The plop project is part a new approach to learning programs being developed at Google and elsewhere that takes on the challenges of learning programs through a unified approach based on reducing programs to a hierarchical normal form, building sequences of specialized representations for programs as search progresses, maintaining alternative representations, and managing uncertainty probabilistically by applying estimation-of-distribution algorithms over program spaces, and exploiting probabilistic background knowledge.

For more information on this approach to learning programs, see my doctoral dissertation. For more on the overall philosophy and where things are going, see the plop wiki on Google Code.
Read More
Posted in | No comments

Friday, 3 October 2008

New Technology Roundtable Series

Posted on 10:28 by Unknown
Posted by Alfred Spector, VP of Research and Special Initiatives

We've just posted the first three videos in the Google Technology Roundtable Series.  Each one is a discussion with senior Google researchers and technologists about one of our most significant achievements. We use a talk show format, where I lead a discussion on the technology. 

While the videos are intended for a reasonably technical audience, I think they may be interesting to many as an overview of the key challenges and ideas underlying Google's systems. And of course they offer a glimpse into the people behind Google.

The first one we made is Large-Scale Search System Infrastructure and Search Quality." I interview Google Fellows Jeff Dean and Amit Singhal on their insights in how search works at Google.

The next title is "Map Reduce," a discussion of this key technology (first, at Google, and now having a great impact across the field) for harnessing parallelism provided by very large-scale clusters computers, while mitigating the component failures that inevitably occur in such big systems. My discussion is with four of our Map Reduce expert engineers: Sanjay Ghemawat and Jeff Dean again, plus Software Engineers Jerry Zhao and Matt Austern, who discuss the origin, evolution and future of Map Reduce. By the way, this type of infrastructure underlies the infrastructure concepts in our recent post on "The Intelligent Cloud."

The third video, "Applications of Human Language Technology," is a discussion of our enormous progress in large-scale automated translation of languages and speech recognition. Both of these technology domains are coming of age with capabilities that will truly impact what we expect of computers on a day-to-day basis. I discuss these technologies with human language technology experts Franz Josef Och, an expert in the automated translation of languages, and Mike Cohen, an expert in speech processing.

We hope to produce more of these, so please leave feedback at YouTube (in the comments field for each video), and we will incorporate your ideas into our future efforts.
Read More
Posted in | No comments

Monday, 29 September 2008

Doubling Up

Posted on 20:06 by Unknown
Posted by Franz Josef Och

Machine translation is hard. Natural languages are so complex and have
so many ambiguities and exceptions that teaching a computer to
translate between them turned out to be a much harder problem than
people thought when the field of machine translation was born over 50
years ago. At Google Research, our approach is to have the machines
learn to translate by using learning algorithms on gigantic amounts of
monolingual and translated data. Another knowledge source is user
suggestions. This approach allows us to constantly improve the
quality of machine translations as we mine more data and
get more and more feedback from users.

A nice property of the learning algorithms that we use is that they
are largely language independent -- we use the same set of core
algorithms for all languages. So this means if we find a lot of
translated data for a new language, we can just run our algorithms and
build a new translation system for that language.

As a result, we were recently able to significantly increase the number of
languages on translate.google.com. Last week, we launched eleven new
languages: Catalan, Filipino, Hebrew, Indonesian, Latvian, Lithuanian, Serbian,
Slovak, Slovenian, Ukrainian, Vietnamese. This increases the
total number of languages from 23 to 34.  Since we offer translation
between any of those languages this increases the number of language
pairs from 506 to 1122 (well, depending on how you count simplified
and traditional Chinese you might get even larger numbers). We're very
happy that we can now provide free online machine translation for many
languages that didn't have any available translation system before.

So how far can we go with adding new languages in the future? Can we
go to 40, 50 or even more languages?  It is certainly getting harder,
as less data is available for those languages and as a result it is
harder to build systems that meet our quality bar.  But we're working
on better learning algorithms and new ways to mine data and so even if
we haven't covered your favorite language yet, we hope that we will have
it soon.
Read More
Posted in | No comments

Saturday, 26 July 2008

Remembering Randy Pausch

Posted on 00:51 by Unknown
Posted by Kevin McCurley, Research Team

It is with great sadness that we note the passing of Randy Pausch, who taught computer science at Carnegie Mellon University. Randy was well-known by many within the research community, including quite a number of us here at Google. Alfred Spector, our Vice President of Research, was his Ph.D. advisor. Rich Gossweiler, a Senior Research Scientist, was his first Ph.D. student. Several other former colleagues and coauthors (Joshua Bloch, Adam Fass, and Ning Hu) now work here.

All of us strive to make an impact with our research, and Randy was no exception. He will be remembered for his work, but also for his contributions to humanity at large. Millions have watched the video on YouTube from his lecture titled Achieving your Childhood Dreams. The strength of his character was already known to his family, his colleagues, and the broader computer science research community. The courage and optimism that he displayed at the end of his life became inspirational to millions more.

I've seen Randy repeatedly go to bat for what is right. As a leader, he consistently evoked incredible enthusiasm and optimism for the subjects he embraces. Randy had a very human passion about people and not just who they are, but their potential, despite any flaws or obstacles in their way. His contributions will be remembered for generations to come. - Rich Gossweiler

Randy was one of the most vibrant, passionate people I've ever known. His passion was inspirational not only to his family and colleagues, but also, because of his courageous presentations beginning with his well-known Last Lecture, he has influenced millions more. - Alfred Spector


We will miss Randy very much, and remember him fondly.
Read More
Posted in | No comments
Newer Posts Older Posts Home
Subscribe to: Posts (Atom)

Popular Posts

  • Our Faculty Institute brings faculty back to the drawing board
    Posted by Nina Kim Schultz, Google Education Research Cross-posted with the Official Google Blog School may still be out for summer, but tea...
  • Academic Successes in Cluster Computing
    Posted by Alfred Spector, VP of Research Access to massive computing resources is foundational to Research and Development. Fifteen awardees...
  • Towards Energy-Proportional Datacenters
    Posted by Dennis Abts, Michael R. Marty, Philip M. Wells, Peter Klausler, and Hong Liu This is part of the series highlighting some notable...
  • International Conference on Machine Learning (ICML 2009) in Montreal
    Posted by Eyal Even Dar and Vahab Mirrokni , Google Research, NY The 26th International Conference on Machine Learning ( ICML 2009 ) was re...
  • A new landmark in computer vision
    Posted by Jay Yagnik, Head of Computer Vision Research [Cross-posted with the Official Google Blog ] Science fiction books and movies have l...
  • Market Algorithms and Optimization Meeting
    Posted by  Vahab S. Mirrokni and Muthu Muthukrishnan Google auctions ads, and enables a market with millions of advertisers and users.  This...
  • Education Awards on Google App Engine
    Posted by Andrea Held, Google University Relations Cross-posted with Google Developers Blog Last year we invited proposals for innovative p...
  • Speed Matters
    Posted by Jake Brutlag, Web Search Infrastructure At Google, we've gathered hard data to reinforce our intuition that "speed matter...
  • Google launches Korean Voice Search
    Posted by Mike Schuster & Martin Jansche, Google Research On June 16th, we launched our Korean voice search system . Google Search by Vo...
  • Two Views from the 2009 Google Faculty Summit
    Posted by Alfred Spector, Vice President of Research and Special Initiatives [cross-posted with the Official Google Blog ] We held our fifth...

Categories

  • accessibility
  • ACL
  • ACM
  • Acoustic Modeling
  • ads
  • adsense
  • adwords
  • Africa
  • Android
  • API
  • App Engine
  • App Inventor
  • Audio
  • Awards
  • Cantonese
  • China
  • Computer Science
  • conference
  • conferences
  • correlate
  • crowd-sourcing
  • CVPR
  • datasets
  • Deep Learning
  • distributed systems
  • Earth Engine
  • economics
  • Education
  • Electronic Commerce and Algorithms
  • EMEA
  • EMNLP
  • entities
  • Exacycle
  • Faculty Institute
  • Faculty Summit
  • Fusion Tables
  • gamification
  • Google Books
  • Google+
  • Government
  • grants
  • HCI
  • Image Annotation
  • Information Retrieval
  • internationalization
  • Interspeech
  • jsm
  • jsm2011
  • K-12
  • Korean
  • Labs
  • localization
  • Machine Hearing
  • Machine Learning
  • Machine Translation
  • MapReduce
  • market algorithms
  • Market Research
  • ML
  • MOOC
  • NAACL
  • Natural Language Processing
  • Networks
  • Ngram
  • NIPS
  • NLP
  • open source
  • operating systems
  • osdi
  • osdi10
  • patents
  • ph.d. fellowship
  • PiLab
  • Policy
  • Public Data Explorer
  • publication
  • Publications
  • renewable energy
  • Research Awards
  • resource optimization
  • Search
  • search ads
  • Security and Privacy
  • SIGMOD
  • Site Reliability Engineering
  • Speech
  • statistics
  • Structured Data
  • Systems
  • Translate
  • trends
  • TV
  • UI
  • University Relations
  • UNIX
  • User Experience
  • video
  • Vision Research
  • Visiting Faculty
  • Visualization
  • Voice Search
  • Wiki
  • wikipedia
  • WWW
  • YouTube

Blog Archive

  • ▼  2013 (51)
    • ▼  December (3)
      • Groundbreaking simulations by Google Exacycle Visi...
      • Googler Moti Yung elected as 2013 ACM Fellow
      • Free Language Lessons for Computers
    • ►  November (9)
    • ►  October (2)
    • ►  September (5)
    • ►  August (2)
    • ►  July (6)
    • ►  June (7)
    • ►  May (5)
    • ►  April (3)
    • ►  March (4)
    • ►  February (4)
    • ►  January (1)
  • ►  2012 (59)
    • ►  December (4)
    • ►  October (4)
    • ►  September (3)
    • ►  August (9)
    • ►  July (9)
    • ►  June (7)
    • ►  May (7)
    • ►  April (2)
    • ►  March (7)
    • ►  February (3)
    • ►  January (4)
  • ►  2011 (51)
    • ►  December (5)
    • ►  November (2)
    • ►  September (3)
    • ►  August (4)
    • ►  July (9)
    • ►  June (6)
    • ►  May (4)
    • ►  April (4)
    • ►  March (5)
    • ►  February (5)
    • ►  January (4)
  • ►  2010 (44)
    • ►  December (7)
    • ►  November (2)
    • ►  October (9)
    • ►  September (7)
    • ►  August (2)
    • ►  July (7)
    • ►  June (3)
    • ►  May (2)
    • ►  April (1)
    • ►  March (1)
    • ►  February (1)
    • ►  January (2)
  • ►  2009 (44)
    • ►  December (8)
    • ►  November (4)
    • ►  August (4)
    • ►  July (5)
    • ►  June (5)
    • ►  May (4)
    • ►  April (6)
    • ►  March (3)
    • ►  February (1)
    • ►  January (4)
  • ►  2008 (11)
    • ►  December (1)
    • ►  November (1)
    • ►  October (1)
    • ►  September (1)
    • ►  July (1)
    • ►  May (3)
    • ►  April (1)
    • ►  March (1)
    • ►  February (1)
  • ►  2007 (9)
    • ►  October (1)
    • ►  September (2)
    • ►  August (1)
    • ►  July (1)
    • ►  June (2)
    • ►  February (2)
  • ►  2006 (15)
    • ►  December (1)
    • ►  November (1)
    • ►  September (1)
    • ►  August (1)
    • ►  July (1)
    • ►  June (2)
    • ►  April (3)
    • ►  March (4)
    • ►  February (1)
Powered by Blogger.

About Me

Unknown
View my complete profile