Monday, 31 August 2015

VLDB 2015 and Database Research at Google



This week, Kohala, Hawaii hosts the 41st International Conference of Very Large Databases (VLDB 2015), a premier annual international forum for data management and database researchers, vendors, practitioners, application developers and users. As a leader in Database research, Google will have a strong presence at VLDB 2015 with many Googlers publishing work, organizing workshops and presenting demos.

The research Google is presenting at VLDB involves the work of Structured Data teams who are building intelligent and efficient systems to discover, annotate and explore structured data from the Web, surfacing them creatively through Google products (such as structured snippets and table search), as well as engineering efforts that create scalable, reliable, fast and general-purpose infrastructure for large-scale data processing (such as F1, Mesa, and Google Cloud's BigQuery).

If you are attending VLDB 2015, we hope you’ll stop by our booth and chat with our researchers about the projects and opportunities at Google that go into solving interesting problems for billions of people. You can also learn more about our research being presented at VLDB 2015 in the list below (Googlers highlighted in blue).

Google is a Gold Sponsor of VLDB 2015.

Papers:
Keys for Graphs
Wenfei Fan, Zhe Fan, Chao Tian, Xin Luna Dong

In-Memory Performance for Big Data
Goetz Graefe, Haris Volos, Hideaki Kimura, Harumi Kuno, Joseph Tucek, Mark Lillibridge, Alistair Veitch

The Dataflow Model: A Practical Approach to Balancing Correctness, Latency, and Cost in Massive-Scale, Unbounded, Out-of-Order Data Processing
Tyler Akidau, Robert Bradshaw, Craig Chambers, Slava Chernyak, Rafael Fernández-Moctezuma, Reuven Lax, Sam McVeety, Daniel Mills, Frances Perry, Eric Schmidt, Sam Whittle

Resource Bricolage for Parallel Database Systems
Jiexing Li, Jeffrey Naughton, Rimma Nehme

AsterixDB: A Scalable, Open Source BDMS
Sattam Alsubaiee, Yasser Altowim, Hotham Altwaijry, Alex Behm, Vinayak Borkar, Yingyi Bu, Michael Carey, Inci Cetindil, Madhusudan Cheelangi, Khurram Faraaz, Eugenia Gabrielova, Raman Grover, Zachary Heilbron, Young-Seok Kim, Chen Li, Guangqiang Li, Ji Mahn Ok, Nicola Onose, Pouria Pirzadeh, Vassilis Tsotras, Rares Vernica, Jian Wen, Till Westmann

Knowledge-Based Trust: A Method to Estimate the Trustworthiness of Web Sources
Xin Luna Dong, Evgeniy Gabrilovich, Kevin Murphy, Van Dang, Wilko Horn, Camillo Lugaresi, Shaohua Sun, Wei Zhang

Efficient Evaluation of Object-Centric Exploration Queries for Visualization
You Wu, Boulos Harb, Jun Yang, Cong Yu

Interpretable and Informative Explanations of Outcomes
Kareem El Gebaly, Parag Agrawal, Lukasz Golab, Flip Korn, Divesh Srivastava

Take me to your leader! Online Optimization of Distributed Storage Configurations
Artyom Sharov, Alexander Shraer, Arif Merchant, Murray Stokely

TreeScope: Finding Structural Anomalies In Semi-Structured Data
Shanshan Ying, Flip Korn, Barna Saha, Divesh Srivastava

Workshops:
Workshop on Big-Graphs Online Querying - Big-O(Q) 2015
Workshop co-chair: Cong Yu

3rd International Workshop on In-Memory Data Management and Analytics
Program committee includes: Sandeep Tata

High-Availability at Massive Scale: Building Google's Data Infrastructure for Ads
Invited talk at BIRTE by: Ashish Gupta, Jeff Shute

Demonstrations:
KATARA: Reliable Data Cleaning with Knowledge Bases and Crowdsourcing
Xu Chu, John Morcos, Ihab Ilyas, Mourad Ouzzani, Paolo Papotti, Nan Tang, Yin Ye

Error Diagnosis and Data Profiling with Data X-Ray
Xiaolan Wang, Mary Feng, Yue Wang, Xin Luna Dong, Alexandra Meliou

Friday, 28 August 2015

Announcing Google’s 2015 Global PhD Fellows



In 2009, Google created the PhD Fellowship program to recognize and support outstanding graduate students doing exceptional research in Computer Science and related disciplines. Now in its seventh year, our fellowship programs have collectively supported over 200 graduate students in Australia, China and East Asia, India, North America, Europe and the Middle East who seek to shape and influence the future of technology.

Reflecting our continuing commitment to building mutually beneficial relationships with the academic community, we are excited to announce the 44 students from around the globe who are recipients of the award. We offer our sincere congratulations to Google’s 2015 Class of PhD Fellows!

Australia

  • Bahar Salehi, Natural Language Processing (University of Melbourne)
  • Siqi Liu, Computational Neuroscience (University of Sydney)
  • Qian Ge, Systems (University of New South Wales)

China and East Asia

  • Bo Xin, Artificial Intelligence (Peking University)
  • Xingyu Zeng, Computer Vision (The Chinese University of Hong Kong)
  • Suining He, Mobile Computing (The Hong Kong University of Science and Technology)
  • Zhenzhe Zheng, Mobile Networking (Shanghai Jiao Tong University)
  • Jinpeng Wang, Natural Language Processing (Peking University)
  • Zijia Lin, Search and Information Retrieval (Tsinghua University)
  • Shinae Woo, Networking and Distributed Systems (Korea Advanced Institute of Science and Technology)
  • Jungdam Won, Robotics (Seoul National University)

India

  • Palash Dey, Algorithms (Indian Institute of Science)
  • Avisek Lahiri, Machine Perception (Indian Institute of Technology Kharagpur)
  • Malavika Samak, Programming Languages and Software Engineering (Indian Institute of Science)

Europe and the Middle East

  • Heike Adel, Natural Language Processing (University of Munich)
  • Thang Bui, Speech Technology (University of Cambridge)
  • Victoria Caparrós Cabezas, Distributed Systems (ETH Zurich)
  • Nadav Cohen, Machine Learning (The Hebrew University of Jerusalem)
  • Josip Djolonga, Probabilistic Inference (ETH Zurich)
  • Jakob Julian Engel, Computer Vision (Technische Universität München)
  • Nikola Gvozdiev, Computer Networking (University College London)
  • Felix Hill, Language Understanding (University of Cambridge)
  • Durk Kingma, Deep Learning (University of Amsterdam)
  • Massimo Nicosia, Statistical Natural Language Processing (University of Trento)
  • George Prekas, Operating Systems (École Polytechnique Fédérale de Lausanne)
  • Roman Prutkin, Graph Algorithms (Karlsruhe Institute of Technology)
  • Siva Reddy, Multilingual Semantic Parsing (The University of Edinburgh)
  • Immanuel Trummer, Structured Data Analysis (École Polytechnique Fédérale de Lausanne)
  • Margarita Vald, Security (Tel Aviv University)

North America

  • Waleed Ammar, Natural Language Processing (Carnegie Mellon University)
  • Justin Meza, Systems Reliability (Carnegie Mellon University)
  • Nick Arnosti, Market Algorithms (Stanford University)
  • Osbert Bastani, Programming Languages (Stanford University)
  • Saurabh Gupta, Computer Vision (University of California, Berkeley)
  • Masoud Moshref Javadi, Computer Networking (University of Southern California)
  • Muhammad Naveed, Security (University of Illinois at Urbana-Champaign)
  • Aaron Parks, Mobile Networking (University of Washington)
  • Kyle Rector, Human Computer Interaction (University of Washington)
  • Riley Spahn, Privacy (Columbia University)
  • Yun Teng, Computer Graphics (University of California, Santa Barbara)
  • Carl Vondrick, Machine Perception, (Massachusetts Institute of Technology)
  • Xiaolan Wang, Structured Data (University of Massachusetts Amherst)
  • Tan Zhang, Mobile Systems (University of Wisconsin-Madison)
  • Wojciech Zaremba, Machine Learning (New York University)

Thursday, 27 August 2015

Real-Time Data Validation with Google Tag Assistant Recordings

We’ve said it before and we’ll say it again: great analytics can only happen with great data.  

That's why we've made it a priority to help our users confirm that their data is top-quality. Last year we released our automated data diagnostics feature, and now we’re proud to announce the launch of another powerful new feature: Google Tag Assistant Recordings.  

This tool helps you instantly validate your Google Analytics or Google Analytics Premium implementation. If it finds data quality issues, it helps you troubleshoot them and then recheck them on the spot.  It’s available as part of the Google Tag Assistant Chrome Extension.
Screen Shot 2015-07-21 at 2.27.31 PM.png
"Tag Assistant Recordings is fast becoming one of my favorite tools for debugging Google Analytics Premium installations!  I use it multiple times a day with my Premium clients to help explain odd trends in their data or debug configuration issues. Already I'm building it into my core workflow." 

- Dan Rowe, Director of Analytics at Analytics Pros

What can I use it for?

Tag Assistant Recordings works with all kinds of data events: purchases, logins, and so on. What if you sell flowers online and want to confirm that Enhanced Ecommerce is capturing the checkout flow correctly? With Tag Assistant Recordings, you can record yourself going through the checkout process as you buy a dozen red roses, and then review what Google Analytics captured.

If you find that your account isn’t set up properly — if the sale wasn't recorded or was mis-labeled — you can make adjustments and test it all over again instantly.  With Tag Assistant Recordings, you know you’re capturing all the data that’s important to you.

Tag Assistant Recordings can be particularly useful when (1) you’re in the process of implementing Google Analytics or Google Analytics Premium, (2) you’ve recently made updates to your site, or (3) you’re making changes to your Google Analytics or Google Analytics Premium configuration. It works even if your new site or your updates aren't visible to the public yet, so you can feel confident before you go live.

Tag Assistant Recordings can also help if you want to reconfigure your Google Analytics account to better reflect your business.  For example, you may want to configure multi-channel funnels to detect your AdWords channel.  Tag Assistant Recordings lets you set up this new functionality in Google Analytics and test immediately whether everything is working as you expect.  

"Tag Assistant Recordings has already been a HUGE help! Analytics Pros and About.com were working on an issue with sessions double-counting and Tag Assistant Recordings let us narrow down precisely which hits were having new sessions counted. It saved us hours of time and helped us jump right to where the problem was. So, in summary, this is awesome!"  

- Greg McDonald, Business Intelligence Analyst at About.com

How does it work?

Tag Assistant Recordings works through the Google Tag Assistant Chrome Extension, so you’ll need to download the extension if you aren’t already using it.  From there, setup is easy.  Simply open Google Tag Assistant, record the user flow you’d like to check, and then view the full report in Tag Assistant.  You’ll want to view both tabs in the report (Tag Assistant and Google Analytics) to verify that you see the intended tags.  Keep in mind that the Google Analytics data is only available if you have access to the appropriate property or view.

Tag Recordings Gif.gif

Here's a nifty bonus: If you find a problem, and you think you have fixed it by changing settings from within Google Analytics, return to the Google Analytics tab in Tag Assistant Recordings and click the “Update” button. You'll see instantly how your configuration changes would have affected this recording.

We hope that Google Tag Assistant will be a valuable new tool in your analytics toolkit.  

Why not start using it today?


Posted by:  Ajay Nainani, Frank Kieviet, and Jocelyn Whittenburg, Google Analytics team

Friday, 21 August 2015

Affiliate Attribution: Putting the Pieces Together

Originally Posted on the Adometry M2R Blog
Recently I was reminded of an article from a little while back, titled, “2013: The Year of Affiliate Attribution?” It’s an interesting take and worthwhile read for those interested in affiliate marketing and the associated measurement challenges. Given that some time has passed, I thought it would be interesting to take a look at progress to date towards realizing a more holistic and accurate view of affiliate performance as part of a comprehensive cross-channel strategy.
Most affiliate managers have a similar goal to manage affiliate holistically, meaning investing in those that predominantly drive net-new customers independent of other paid marketing investments. Ultimately, this model allows them to optimize CPA by managing commissions, coupon discounts, and brand appropriateness based on true “incremental value” provided to business. Unfortunately, due to a lack of transparency and inadequate measurement, many marketers find themselves short of this goal. The result is the ongoing nagging question, “Is my affiliate strategy working and am I overpaying for what I’m getting?”

Why ‘Affiliate Attribution’ Is Hard

Affiliate marketers’ challenges range from competing against affiliates in PPC ad programs to concerns about questionable business practices employed by some “opportunistic” affiliates offering marginal value, but still receiving credit for sales that likely would have happened regardless. Which brings us to the central question:
How do marketers determine how much credit an affiliate should receive?

As you may know, opinions about how much conversion credit affiliates deserve for any given transaction vary widely. While there are a number of factors that influence affiliate performance (e.g. where they appear in the sales funnel, industry/sector, time-to-purchase length, etc.) for most brands the attribution model that is utilized will have a significant impact on which affiliates are over- and under-valued.
For example, in a last-click world affiliates that enter the purchase path towards the bottom of the funnel often hold their own; yet, when brands begin measuring on a full-funnel basis incorporating impression data, many struggle to prove their incremental value as the consumer has many exposures to marketing long before they reach the affiliate site. Conversely, affiliates that act predominantly as top- or mid-funnel (content, loyalty, etc.) are usually undervalued using last-click but can garner more credit using a full-funnel, data-driven attribution methodology. I should also mention these are broad generalizations only meant as examples, and it’s not necessarily a zero-sum game.
Another challenge is that fractional, data-driven attribution is difficult to implement for some types of promotions. One instance of this is cash back, loyalty and reward sites that must know an exact commission amount they will receive for each transaction so that they can pass on discounts to members. Given the complexity of more sophisticated attribution models, this data isn’t readily available.
Lastly, there several organizational challenges that inhibit the use of data-driven attribution among affiliate marketers. Some industry experts have indicated that many publishers, as much as 70-80%, strip impression tracking code from affiliate URLs. Another measurement challenge we see frequently is brands managing affiliates at the channel level leaving little sub-channel categorization which is where significant optimization opportunities exist.
Affiliate Attribution and the Performance Marketing Goldmine
Of course, part of our work at Adometry is helping customers address these challenges (and more) to ensure they are measuring affiliate contributions accurately and able to take appropriate action based on fully-attributed results.
Some key advantages of using data-driven attribution to measure affiliate sales include:
  • The ability to create a unified framework to compare performance (clicks and Impressions) in which affiliates compete for budgets on equal footing,
  • Increased visibility into which publishers are truly driving net-new customers through specifying which are an integral part of a multi-touch path and which are expendable,
  • The knowledge required to implement a Publisher category taxonomy to allow more insights into how different types of publishers perform by funnel stage and areas to improve efficiency,
  • Insight into the true incremental value publishers are providing and the offering commission rates to reflect this actual value,
  • A better understanding of affiliate’s role in the overall mix, further informing marketers use of complementary tactics to maximize affiliate contributions in concert with other channels,
  • The ability to use actual performance data to counter myths and frustrations with affiliates (cookie stuffing, stealing conversions, etc.)
Taken separately, each of these represents a significant opportunity to both be more effective in how you identify and utilize affiliate attribution to drive new opportunities. Together, they represent a fundamental improvement in how you manage your overall marketing spending, strategic planning and optimization efforts.
Top-performing affiliates, particularly those at the top and middle of the funnel, also stand to benefit from more transparent, accurate and fair system for crediting conversions. In fact, several large-scale, forward-thinking affiliates are already investing in data-driven attribution to arm themselves with the data required to effectively compete and win business in the market as brands become more sophisticated and judicious with their affiliates budgets.
It’s an exciting time for performance marketing. Change is always hard, but in this case it’s absolutely change for the better.  And frankly, its time.  What are your thoughts and experiences with measuring affiliate performance and attribution?

Posted by Casey Carey, Google Analytics team

Google Faculty Research Awards: Summer 2015



We have just completed another round of the Google Faculty Research Awards, our annual open call for research proposals on Computer Science and related topics, including systems, machine learning, software engineering, security and mobile. Our grants cover tuition for a graduate student and provide both faculty and students the opportunity to work directly with Google researchers and engineers.

This round we received 805 proposals, about the same as last round, covering 48 countries on 6 continents. After expert reviews and committee discussions, we decided to fund 113 projects, with 27% of the funding awarded to universities outside the U.S. The subject areas that received the highest level of support were systems, machine perception, software engineering, and machine learning.

The Faculty Research Awards program plays a critical role in building and maintaining strong collaborations with top research faculty globally. These relationships allow us to keep a pulse on what’s happening in academia in strategic areas, and they help to extend our research capabilities and programs. Faculty also report, through our annual survey, that they and their students benefit from a direct connection to Google as a source of ideas and perspective.

Congratulations to the well-deserving recipients of this round’s awards. If you are interested in applying for the next round (deadline is October 15), please visit our website for more information.

Thursday, 20 August 2015

The Next Chapter for Flu Trends



When a small team of software engineers first started working on Flu Trends in 2008, we wanted to explore how real-world phenomena could be modeled using patterns in search queries. Since its launch, Google Flu Trends has provided useful insights and served as one of the early examples for “nowcasting” based on search trends, which is increasingly used in health, economics, and other fields. Over time, we’ve used search signals to create prediction models, updating and improving those models over time as we compared our prediction to real-world cases of flu.

Instead of maintaining our own website going forward, we’re now going to empower institutions who specialize in infectious disease research to use the data to build their own models. Starting this season, we’ll provide Flu and Dengue signal data directly to partners including Columbia University’s Mailman School of Public Health (to update their dashboard), Boston Children’s Hospital/Harvard, and Centers for Disease Control and Prevention (CDC) Influenza Division. We will also continue to make historical Flu and Dengue estimate data available for anyone to see and analyze.

Flu continues to affect millions of people every year, and while it’s still early days for nowcasting and similar tools for understanding the spread of diseases like flu and dengue fever—we’re excited to see what comes next. To download the historical data or learn more about becoming a research partner, please visit the Flu Trends web page.

Wednesday, 19 August 2015

Google Analytics User Conference: G’day Australia

The Australian Google Analytics User Conference is worth clearing your diaries for, with some of the most well-known and respected international industry influencers making their way to Sydney and Melbourne to present at the conference this September.

Hosted by Google Certified Partners, Loves Data, you’ll be learning about the latest features, what’s trending and popular, best practices and uncovering ways to get the most out of Google Analytics. Topics covered include: making sure digital analytics is indispensable to your organisation; applying analytics frameworks to your whole organisation; improving your data quality and collection; data insights you can action; and presenting data to get results.

Presenting the keynote is Jim Sterne, Chairman of the Digital Analytics Association, founder of eMetrics and also known as the godfather of analytics. Joining him are two speakers from Google in the US: Krista Seiden, Google Product Manager and Analytics Advocate and Mike Kwong, Senior Staff Software Engineer.

Other leading international industry influencers presenting at the conference include Simo Ahava (Google Developer Expert; Reaktor), Chris Chapo (Enjoy), Benjamin Mangold (Loves Data), Lea Pica (Consultant, Leapica.com), Chris Samila (Optimizely), Carey Wilkins (Evolytics) and Tim Wilson (Web Analytics Demystified).  

Expect to network with other like-minded data enthusiasts, marketers, developers and strategists, plus get to know the speakers better during the Conference’s Ask Me Anything session. We’ve even covered our bases for those seeking next-level expertise with a marketing or technical masterclass available the day before the conference. Find out more information about the speakers and check out the full program.

Last year’s conference sold out way in advance and this year’s conference is heading in the same direction. Book your tickets now to avoid disappointment. 

Event details Sydney
Masterclass & Conference | 8 & 9 September 2015

Event details Melbourne
Masterclass & Conference | 10 & 11 September 2015

Posted by Will Pryor, Google Analytics team