Due to popular request, here is my overview of some of the coolest stuff from Day 2 of CVPR 2012 in Providence, RI. While the Lobster dinner was the highlight for many of us, there were also some serious learning/optimization-based papers presented during Day 2 worthy of sharing. Here are some of the papers which left me with a very positive impression.
Dennis Strelow of Google Research in Mountain View presented a general framework for Wiberg minimization. This is a strategy for minimizing objective functions with multiple variables -- objectives which are typically tackled in an EM-style fashion. The idea is to express one of the variables as a linear function of the other variable, effectively making the problem depend on only one set of variables. The technique is quite general and has been shown to produce state-of-the-art results on a bundle adjustment problem. I know Dennis from my second internship at Google where we worked on some sparse-coding problems. If you perform lots of matrix decomposition problems, check out his paper!
Dennis Strelow
General and Nested Wiberg Minimization
CVPR 2012
Another cool paper which is all about learning is Hossein Mobahi's algorithm for optimizing objectives by smoothing them to avoiding getting stuck in local minima. This paper is not about blurry images, but about applying Gaussians to objective functions. In fact, for the problem of image alignment, Hossein provides closed form versions of image operators. Now when you apply these operators to images, you efficiently smooth the underlying cross-correlation alignment objective. You decrease the blur, while following the optimum path, and get much nicer answers that doing naive image alignment.
Hossein Mobahi, C. Lawrence Zitnick, Yi Ma
Seeing through the Blur
CVPR 2012
Ira Kemelmacher-Shlizerman, of Photobios fame, showed a really cool algorithm for computing optical flow between two different faces based on learning a subspace (using a large database of faces). The ideas is quite simple and allows for flowing between two very different faces where the underlying operation produces a sequence of intermediate faces in an interpolation-like manner. She shared this video with us during her presentation, but it is on Youtube, so now you can enjoy it for yourself.
Ira Kemelmacher-Shlizerman, Steven M. Seitz
Collection Flow
CVPR 2012
Now talk about cool ideas! Pyry, of CMU fame, presented a recommendation engine for classifiers. The idea is to take techniques from collaborative filtering (think Netflix!) and apply then to the classifier selection problem. Pyry has been working on action recognition and the ideas presented in this work are not only quite general, but have are quite intuitive and likely to benefit anybody working with large collections of classifiers.
Pyry Matikainen, Rahul Sukthankar, Martial Hebert
Model Recommendation for Action Recognition
CVPR 2012
And finally, a super-easy algorithm presented for metric learning by Martin Köstinger had me intrigued! This a Mahalanobis distance metric learning paper which uses equivalence relationships. This means that you are given pairs of similar items and pairs of dissimilar items. The underlying algorithm is really not much more than fitting two covariance matrices, one to the positive equivalence relations, and another to the non-equivalence relations. They have lots of code online, and if you don't believe that such a simple algorithm can beat LMNN (Large-Margin Nearest Neighbor from Killian Weinberger), then get their code and hack away!
Martin Köstinger, Martin Hirzer, Paul Wohlhart, Peter M. Roth, Horst Bischof
Large Scale Metric Learning from Equivalence Constraints
CVPR 2012
CVPR 2012 gave us many very math-oriented papers, and while I cannot list of all of them, I hope you found my short list useful.
you may like
Tampilkan postingan dengan label photobios. Tampilkan semua postingan
Tampilkan postingan dengan label photobios. Tampilkan semua postingan
Rabu, 20 Juni 2012
Rabu, 24 Agustus 2011
The vision hacker culture at Google ...
I sometimes get frustrated when developing machine learning algorithms in C++. And since working in object recognition basically means you have to be a machine learning expert, trying something new and exciting in C++ can be extremely painful. I don't miss the C++ heavy workflow for vision projects at Google. C++ is great for building large-scale systems, but not for pioneering object recognition representations. I like to play with pixels and I like to think of everything as matrices. But programming languages, software engineering philosophies, and other coding issues aren't going to be today's topic. Today I want to talk about the one thing that is more valuable that is computers, and that is people. Not just people, but a community of people, and in particular the culture at Google -- in particular, vision@Google.
The people at Google aren't just hackers, they are Jedis when it comes to building great stuff -- and that is why I recommend a Google internship to many of my fellow CMU vision Robograds (fyi, Robograds are CMU Robotics Graduate Students). CMU-ers, like Googlers, like to build stuff. However, CMU-ers are typically younger.
I miss being around the hacker culture at Google.
Image from http://www.rabittooth.com/
What is a software engineering Jedi, you might ask? Tis' one who is not afraid of million cores, one who is not afraid of building something great. While little boys get hurt by the guns 'n knives of C++, Jedi use their tools like ninjas use their swords. You go into Google as a boy, you come out a man. NOTE: I do not recommend going to Google and just toying around in Matlab for 3 months. Build something great, find a Yoda-esque mentor, or at least strive to be a Jedi. There's plenty of time in grad school for Matlab and writing papers. If you get a chance to go to Google, take the opportunity to go large-scale and learn to MapReduce like the pros.
Every day I learn about more and more people I respect in vision and learning going to Google, or at least interning there (e.g., Andrej Karpathy who is starting his PhD@Stanford and Santosh Divvala who is a well-known CMU PhD student and vision hacker). And I really can't blame them for choosing Google over places like Microsoft for the summer. I can't think of many better places to be -- the culture is inimitable. I spent two summers at Jay Yagnik's group some of the great people I interned with are already full-time Googlers (e.g. Luca Bertelli and Mehmet Emre Sargin). And what is really great about vision@google is that these guys get to publish surprisingly often! Not just throw-away-code kind of publish, but stuff that fits inside large-scale systems -- stuff which is already inside Google products. The technology is often inside the Google product before the paper goes public! Of course it's not easy to publish at a place like Google because there is just way too much exciting large-scale stuff going on. Here is a short list of some cool 2010/2011 vision papers (from vision conferences) with significant Googler contributions.
“Kernelized Structural SVM Learning for Supervised Object Segmentation”, Luca Bertelli, Tianli Yu, Diem Vu, Burak Gokturk, Proceedings of IEEE Conference on Computer Vision and Pattern Recognition 2011.
[abstract] [pdf]
“Finding Meaning on YouTube: Tag Recommendation and Category Discovery”, George Toderici,Hrishikesh Aradhye, Marius Pasca, Luciano Sbaiz, Jay Yagnik, Computer Vision and Pattern Recognition, 2010.
[abstract] [pdf]
Here is a very exciting and new paper from SIGGRAPH 2011. It is a sort of Visual Memex for faces -- congratulations on this paper, guys! Check out the video below.
Exploring Photobios from Ira Kemelmacher on Vimeo
Ira Kemelmacher-Shlizerman, Eli Shechtman, Rahul Garg, Steven M. Seitz. "Exploring Photobios." ACM Transactions on Graphics (SIGGRAPH), Aug 2011. [pdf]
Finally, here is a very mathematical paper with a sexy title from the vision@google team. It will be presented at the upcoming ICCV 2011 Conference in Barcelona -- the same conference where I'll be presenting my Exemplar-SVM paper.
The Power of Comparative Reasoning
Jay Yagnik, Dennis Strelow, David Ross, Ruei-Sung Lin. ICCV 2011. [PDF]
Kernelized Structural SVM Learning
“Kernelized Structural SVM Learning for Supervised Object Segmentation”, Luca Bertelli, Tianli Yu, Diem Vu, Burak Gokturk, Proceedings of IEEE Conference on Computer Vision and Pattern Recognition 2011.
[abstract] [pdf]
Finding Meaning on YouTube
“Finding Meaning on YouTube: Tag Recommendation and Category Discovery”, George Toderici,Hrishikesh Aradhye, Marius Pasca, Luciano Sbaiz, Jay Yagnik, Computer Vision and Pattern Recognition, 2010.
[abstract] [pdf]
Here is a very exciting and new paper from SIGGRAPH 2011. It is a sort of Visual Memex for faces -- congratulations on this paper, guys! Check out the video below.
Exploring Photobios Movie
Exploring Photobios from Ira Kemelmacher on Vimeo
Ira Kemelmacher-Shlizerman, Eli Shechtman, Rahul Garg, Steven M. Seitz. "Exploring Photobios." ACM Transactions on Graphics (SIGGRAPH), Aug 2011. [pdf]
Finally, here is a very mathematical paper with a sexy title from the vision@google team. It will be presented at the upcoming ICCV 2011 Conference in Barcelona -- the same conference where I'll be presenting my Exemplar-SVM paper.
The Power of Comparative Reasoning
Jay Yagnik, Dennis Strelow, David Ross, Ruei-Sung Lin. ICCV 2011. [PDF]
P.S. If you're a fellow vision blogger, then come find me in Barcelona@iccv2011 -- we'll go brag a beer.
Label:
c++,
face memex,
facetime,
google,
hackers,
interns,
internship,
jay yagnik,
jedi,
kernels,
machine learning,
MATLAB,
photobios,
picasa,
publishing,
svm,
visual memex,
visualization,
youtube
Langganan:
Postingan (Atom)



