Tuesday, July 30, 2024

Benedict's Newsletter: No. 551

NO. 551   FREE EDITION   TUE 30 JUL 2024
SPONSORED BY TOKEN2049
Join 20,000 attendees at TOKEN2049, the world's largest Web3 and blockchain event, on 18-19 September. Industry leaders and global cultural icons will converge in Asia's leading economic hub to unveil what lies ahead.

Secure your pass for a festival experience unlike any other: Use code BEN2049 at CHECKOUT now to get 15% off.

My Work

The AI summer

Hundreds of millions of people have tried ChatGPT, but most of them haven't been back. Every big company has done a pilot, but far fewer are in deployment. Some of this is just a matter of time. But LLMs might also be a trap: they look like products and they look magic, but they aren't. Maybe we have to go through the slow, boring hunt for product-market fit after all. LINK

The VR winter continues

Meta has spent at least $50bn on VR and AR so far, but we're still in the VR winter: the devices aren't good enough or cheap enough and the user base is flat. But no matter how good the devices get, how many people will care? LINK

Upgrade to Premium
You're getting the Free edition. Subscribers to the Premium edition got this two days ago on Sunday evening, together with an exclusive column, complete access to the archive of over 500 issues, and more.

News

Meta's Llama 3 beats OpenAI

Meta released the largest version of its newest model, Llama 3.1, and on most benchmarks this is at least as good as the latest and best from OpenAI. And it's open source (sort of) - you can download it yourself today. See this week's column. LINK

Google gives up on killing cookies 

Google has been talking about replacing cross-site cookies in Chrome since at least 2019. As the dominant player in online ads and the maker of the dominant web browser (excluding iOS), it's tried to pull the rest of the industry around some kind of anonymised interest-based targeting, calling this a 'privacy sandbox'. Chrome would know that you've been looking at lots of websites about cars and could tell other websites 'load a car ad on this page' but the actual tracking data would never leave your device. And, meanwhile, Chrome would stop allowing cross-site (AKA third party) cookies by default (Apple's Safari already did this), forcing everyone to switch. 

However, Google has never been able to solve the stake-holder alignment (i.e cat-herding) of persuading privacy regulators, competition regulators, publishers and ad-tech companies that this is a: a good idea and b: would work. Now it looks like Google is giving up: instead of killing 3P cookies, it will keep them, and "introduce a new experience in Chrome that lets people make an informed choice". The devil is in the wording, but that sounds like an 'ask people if they want to block cookies' button. Now the ad industry is scrambling to work out what that means, and what comes next. LINK

OpenAI (finally) does search

OpenAI has finally announced that it will do a search engine powered by ChatGPT. In principle, you can point an LLM at an index of the web and get it to answer questions about that index, or summarise results that a more conventional search engine finds from that index, and perhaps other approaches as well. Bing Copilot, Google 'Search Summaries', Perplexity and a few others are all pursuing versions of this, and now, so is OpenAI. 

I think there are two views on this concept. On one hand, as we saw in the (very brief) excitement about Bing Copilot, and also in the buzz around Perplexity, there is at a minimum a class of search query where the 'job to be done' might be better served by an answer or a summary, not a link to a web page, and so LLMs could displace classic Google search (even if this ends up being done best by Google itself).

But on the other, the error rate inherent to LLMs is an unsolved and quite possibly unsolvable problem: these systems tell you what the answer would probably look like, not what it is, and that may matter too much for too much general web search: this might work best or only work in narrower verticals. We don't know, but, OpenAI will join the effort to find out. LINK

OpenAI burn rates

Accuracy and ranking is one barrier to entry in search - another is just how much money you have, and the Information reports that OpenAI is already on track to burn $5bn this year. I doubt that it will have trouble raising more (and trying to get a share of the firehose of cash that comes from Google search might help), but it's still a reminder that LLMs have unprecedented capital-intensity, especially for a technology that has yet to find broad product-market fit. LINK

LLMs and IPR

Perplexity is in even more trouble with publishers, with Conde Nast now sending a cease-and-desist for its AI 'summaries' that are-perhaps-too-often just copies of other people's work. Meanwhile, someone leaked an internal spreadsheet of training data for Runway, a very buzzy-in-Hollywood generative AI video maker: it appears to have scraped thousands of influencer videos from YouTube, against ToS. 

There is a growing collision between the philosophical view in many AI circles that training-by-looking is no different to what people do (after all, these systems aren't Napster - they can't generally reproduce what's in the training data) and the legal status of 'using' people's property in an entirely new way but without any new model for permission. PERPLEXITY, RUNWAY

Deepmind does maths

Google's DeepMind built systems that solved four of six problems in the International Maths Olympiad (it looks like DeepMind is still doing pure research even as it was re-orged to be more focused on product). Projects like this are valuable in their own right as pure research, but they're also aimed at getting models to be better at 'reasoning', as opposed to 'pattern-matching' (both crude terms). LINK

Social sextortion

There have been a bunch of stories about scammers extorting minors on social platforms, with a few suicides, and now Meta has removed 63k Instagram accounts linked to the so-called 'Yahoo Boys' scammer scene in Nigeria. LINKREPORTS

Meanwhile, The Information reports that Snap has many of the same problems (obviously, because all social messaging platforms do), but, with its smaller scale, struggles to resource the teams that try to address this. (A few years ago Alex Stamos, former Meta CISO, suggested that Meta should offer content moderation as a service to other smaller social platforms.) LINK

Remember Cameo?

This might be one to put in the 'Covid rotation' file with Clubhouse - Cameo was briefly valued at $1bn but is now too hard-up to pay a $600k fine. Sad, but perhaps not surprising. LINK

Apple Maps on the web

Over a decade after the famously-disastrous launch, Apple Maps is now pretty good, but Apple never made it a website (fitting the general apps-are-better approach), but now, suddenly and somewhat randomly, there is a beta that works in a browser. I wonder if this is a regulatory thing? LINK

About

What matters in tech? What's going on, what might it mean, and what will happen next?

I've spent 20 years analysing mobile, media and technology, and worked in equity research, strategy, consulting and venture capital. I'm now an independent analyst. Mostly, that means trying to work out what questions to ask.

Ideas

This week's viral machine learning paper: LLMs collapse when trained 'indiscriminately' on data produced by LLMs. This speaks to the 'model collapse' problem, but needs to be read with caution, since the word 'indiscriminately' is important: this study is based on training that onlyused data output from another model, which is more a proof-of-concept than a realistic scenario. In other words, we can use 'synthetic data', but only in some domains, to some degree, with caution. LINK

More generally, and linked as an example, Alexis Gallagher wrote a useful discussion of what LLMs might be doing, and whether they are reasoning or pattern-matching. It's important to remember that we really don't have a good theoretical model of why LLMs produce such good results, and hence of what would change if we scaled them, used more synthetic data, or anything else. LINK

Bloomberg reports that Apple has spent more than $20bn on TV shows and movies for AppleTV+ without getting much of an audience. And, apparently, it has a reputation for letting the talent spend whatever they want. It remains hard to see why this exists, except as (very expensive) marketing. LINK

Ofcom, the UK TMT regulator, released a report on how many people have seen deep fakes, with some policy ideas. LINK

The WSJ reports that Amazon's Alexa had $25bn of operating losses from 2017-2021, without delivering much tangible business benefit. As I wrote a few years ago, this thing has product-market-fit for consumers as a voice-activated radio/timer/light-switch, but it does not have product-market fit for Amazon (and that's before missing LLMs). LINK, MY ESSAY

A North Korean hacker got a job as a remote IT worker with a security firm. LINK

Outside interests

Apparently, 'motor-doping' is a problem in cycle racing. LINK

Morgan Stanley adds to the endless argument about stock buybacks. LINK

A scientific classification of ravioli. LINK

Data

E-commerce in South-East Asia. LINK

Another top-down macro attempt to quantify the potential impact of LLMs on employment. These exercises always puzzle me in principle: could you have done this for 'smartphones' or 'the web' 20 or 30 years ago? Would you have predicted the impact of cellular networks on taxi drivers? That the web meant newspaper employment would collapse? LINK

Hasbro reported that Monopoly GO, a smartphone game using its (licensed) IP, has grossed over $3bn since launching last April (see the transcript). LINK

Preview from the Premium edition

Llama  

Last spring, an internal memo from Google leaked that said that Google had no competitive advantage in LLMs, and neither did OpenAI. That's been borne out entirely by what's happened since. There are now at least four 'frontier' models that have equivalent performance to the latest and best from OpenAI, including, now, a version of Meta's Llama, and dozens more on any scatter plot of benchmarks. 

In particular, there is also a cluster of models (including, belatedly, one from OpenAI), that have 75-95% of the performance for 5-10% of the inference cost: there is a very very steep curve from 'very good' to 'best' right now. Welcome to the 'feeds and speeds' phase of the market. 

As you can see if you read the long and very detailed technical paper that Meta released for Llama 3.1, making a frontier model requires a lot of engineering expertise, but it does not seem to have fundamental 

 

THIS IS A PREVIEW FROM THE PREMIUM EDITION - PREMIUM SUBSCRIBERS GET THE COMPLETE COLUMN EVERY WEEK. YOU SHOULD UPGRADE.
Upgrade to Premium
You're getting the Free edition. Subscribers to the Premium edition got this two days ago on Sunday evening, together with an exclusive column, complete access to the archive of over 500 issues, and more.
 

No comments:

Post a Comment

US Clears Iran Mines, Your Lifetime Tax Bill, and a Doggie Drive-Thru

The U.S. military has cleared the Iranian sea mines from the Strait of Hormuz's international shipping lanes in what CENTCOM Com...