Showing posts with label Eric_Jang. Show all posts
Showing posts with label Eric_Jang. Show all posts

Thursday, May 23, 2024

What is Eric Jang up to? All Roads Lead to Robotics

Eric Jang (1X Technologies, formerly Halodi Robotics), from his blog post of Mar. 3, All Roads Lead to Robotics. He links to this video:

He comments:

Because we take an end-to-end neural network approach to autonomy, our capability scaling is no longer constrained by how fast we can write code. All of the capabilities in this video involved no coding, it was just learned from data collected and trained on by our Android Operations team.

He also remarks:

1X is the first robotics company (to my knowledge) to have our data collectors train the capabilities themselves. This really decreases the time-to-a-good-model, because the people collecting data can get very fast feedback on how good their data is and how much data they actually need to solve the robotic task. I predict this will become a widespread paradigm in how robot data is collected in the future.

The main substance of his post is entitled "All AI Software Converges to Robotics Software." Why would/might that be so?

ML deployed in a pure software environment is easier because the world of bits is predictable. You can move some bits from A to B and trust that they show up at their destination with perfect integrity. You can make an API call to some server over the Internet and assume that it will just work. Even if it fails, the set of failure modes are known ahead of time so you can handle all of them.

In robotics, all of the information outside of the robot is unknown. Your future sensor observations, given your actions, are unknown. You also don’t know where you are, where anything else is, what will happen if you make contact with something, whether the light turned on after you flipped the switch, or whether you even flipped the switch at all. Even trivial things like telling the difference between riding an elevator down vs. being hoisted up in a gantry is hard, as the forces experienced by the inertial measurement unit (IMU) sensor look similar in both scenarios. A little bit of ignorance propagates very quickly, and soon your robot ends up on the floor having a seizure because it thinks that it still has a chance at maintaining balance.

As our AI software systems start to touch the real world, like doing customer support or ordering your Uber for you, they will run into many of the same engineering challenges that robotics faces today; the longer a program interacts with a source of entropy, the less formal guarantees we can make about the correctness of our program’s behavior. Even if you are not building a physical robot, your codebase ends up looking a lot like a modern robotics software stack. I spend an unreasonable amount of my time implementing more scalable data loaders and logging infrastructure, and making sure that when I log data, I can re-order all of them into a temporally causal sequence for a transformer. Sound familiar? [...]

If you accept the premise that the engineering and infrastructure problems in LLMs are the same as those in robotics, then we should expect that disembodied AGI and robotic AGI happen at roughly the same time. The hardware is ready and all of the pieces are already there in the form of research papers published over the last 10 years.

I've been having similar thoughts, though not with respect to robotics. Rather, I've been thinking about the role of symbolic computing in robust and flexible systems. Given that the world is full of so-called edge cases, at least some of them very important and fruitful, the problems of LLMs will not be solved through add-ons that provide various symbolic capacities, no matter how clever. In the end, it is going to be necessary to re-construct the LLM with symbolic means, and that will prove to be an unending task, as latent space is ever-evolving.

There's more in the post, but this remark stuck out at me:

Any startup that raised 10-100M USD to train their own big neural network from scratch in the last 2 years ended up paying an enormous capex cost for something that basically every AI startup gets for free today. [...] As such, I think the vast majority of successful startups will be the ones that can nimbly ride the tide of open-source weights.

Saturday, January 6, 2024

Eric Jang: AI is Good For You [interview at The Gradient]

The Gradient has an interesting interview with Eric Jang, a roboticist formerly with Google, now Vice President of AI, 1X Technologies.

* (00:00) Intro
* (01:25) Updates since Eric’s last interview
* (06:07) The problem space of humanoid robots
* (08:42) Motivations for the book “AI is Good for You”
* (12:20) Definitions of AGI
* (14:35) ~ AGI timelines ~
* (16:33) Do we have the ingredients for AGI?
* (18:58) Rediscovering old ideas in AI and robotics
* (22:13) Ingredients for AGI
* (22:13) Artificial Life
* (25:02) Selection at different levels of information—intelligence at different scales
* (32:34) AGI as a collective intelligence
* (34:53) Human in the loop learning
* (37:38) From getting correct answers to doing things correctly
* (40:20) Levels of abstraction for modeling decision-making — the neurobiological stack
* (44:22) Implementing loneliness and other details for AGI
* (47:31) Experience in AI systems
* (48:46) Asking for Generalization
* (49:25) Linguistic relativity
* (52:17) Language vs. complex thought and Fedorenko experiments
* (54:23) Efficiency in neural design
* (57:20) Generality in the human brain and evolutionary hypotheses
* (59:46) Embodiment and real-world robotics
* (1:00:10) Moravec’s Paradox and the importance of embodiment
* (1:05:33) How embodiment fits into the picture—in verification vs. in learning
* (1:10:45) Nonverbal information for training intelligent systems
* (1:11:55) AGI and humanity
* (1:12:20) The positive future with AGI
* (1:14:55) The negative future — technology as a lever
* (1:16:22) AI in the military
* (1:20:30) How AI might contribute to art
* (1:25:41) Eric’s own work and a positive future for AI
* (1:29:27) Outro

Links:

* Eric’s book (https://evjang.com/book/)
* Eric’s Twitter (https://x.com/ericjang11?s=20) and homepage (https://evjang.com/)

Eric's final comments: What the future holds

One example I'd like to like kind of say here is like, it's not really about taking our labor supply and then, you know, swapping it out with with robots.

It's more about like, how can we create a world where there is 10x more labor? And I think people today don't have an answer like so people who are afraid of robots taking over the jobs and such. They don't want their own jobs replaced, but they also don't have an answer to as to how we can 10x the volume of labor supply, right?

And I think if you really frame the question in terms of like, in order to make the world better, You do need more labor. And so the labor pool actually needs to increase. And I guess short of just TEDxing the world population, you do need to just make a bunch of robots to do this. So that's kind of the new vision I have for how my career can fill this.

And as the path to AGI, This is not a direct way to AGI. It's more just like I want to build really, really good systems that can do tasks at a high level of success. And I think this will be a really good stepping stone towards actually building useful AGI systems through the mastery of things like deep learning.

Monday, March 20, 2023

Language, LLMs, and culture: Out of Plato's cave

Jon Evans, Language is out Latent Space, Gradient Ascendant, March 14, 2023.

An analogy for how LLMs work:

Another analogy, as two combined can be more illuminating than one: consider snooker, the pool-like game won by sinking balls of varying value in the best possible order. Imagine a snooker table the size of Central Park, occupied by thousands of pockets and millions of numbered balls (the numbers 1 through 32000, repeated.) Now imagine that the rules of snooker — i.e. which balls are most profitable to sink — change after every shot, depending on where the cue ball is, which balls have previously been sunk, the phase of the moon, etc.

Call that “Jungle Snooker”, borrowing from Eric Jang's idea of Jungle Basketball. The numbers on the balls represent word embeddings; ‘which balls to aim to sink in which order,’ the patterns in latent space. All we have really taught modern LLMs is how to be extremely (stochastically) good at Jungle Snooker, which doesn’t feel that different, qualitatively, from teaching them how to be extremely good at Go or chess. Now, the results, when converted into words, are phenomenal, often eerie —

— but LLMs still don’t “know” that their numbers represent words. In fact they never see words per se; we actually break language into tokens, word fragments basically like phonemes, number those tokens, and feed those numbers in as inputs.

Language as the latent space of culture:

Our latent space, known as language, implicitly encodes an enormous amount of knowledge about the world: concepts, relationships, interactions, constraints. LLM embeddings in turn implicitly include a distilled version of that knowledge. A reason LLMs are so unreasonably effective is that language itself is a machine for understanding, one which, it turns out, includes undocumented and previously unused capabilities — a “capability overhang.”

You might also look at Ted Underwood's paper, Mapping the latent spaces of culture.

Out of the cave:

Invert Plato's cave, and imagine yourself as a puppet master trying to reach out to chained prisoners to whom you can only communicate with shadows. Similarly, right now all we have are machines that we can teach to play Jungle Snooker.

But if we do ever build a machine capable of genuine understanding -- setting aside the question of whether we want to, and noting that people in the field generally think it's “when” not “if” -- it seems likely that language will be our most effective shacklebreaker, just as it was for us. This in turn means today's LLMs are likely to be the crucially important first step down that path.

The question is whether there is any iterative path from Jungle Snooker to Plato's Cave to emergence. Some people think we'll just scale there, and as machines get better at Jungle Snooker, they will naturally develop a facility for abstracting complexity into heuristics, which will breed agency and curiosity and a kind of awareness — or at least behavior indistinguishable from awareness — in the same way that embeddings and latent space spontaneously emerge when you teach LLMs.

Others (including me) suspect that whole new fundamental architectures and/or training techniques will be required. But either way, it seems very likely that language will be key, and that modern LLMs, though they'll seem almost comically crude in even five years, are a historically important technology. Language is our latent space, and that's what gives it its unreasonable power.

Yes, we'll need a whole new architecture. There's more at the link.

Thursday, December 15, 2022

We’ve stepped over the threshold into the Fourth Arena, but don’t recognize it

First there was the world of inanimate matter, the First Arena. Life arose from that, the Second Arena. Only 100s of thousands of years ago humans evolved from higher primates and the Third Arena, human culture, appeared. We are now on the threshold of the Fourth Arena. How do we characterize it?

Note: These thoughts are off the top of my head [thinking-out-loud]. It was all I could do to get them out. That’s enough for now. Refinement will have to wait.

* * * * *

OpenAI released GPT-3 in the summer of 2020. It was clear to me that, yes, it has that potential. I registered my response, initially in a comment over at Marginal Revolution, and then on New Savanna, First thoughts on the implications of GPT-3 [here be dragons, we're swimming and flying with them]. Having explored ChatGPT for the past two weeks, that potential emerges before me with even greater clarity. Of course, we could blow it, nothing is guaranteed. But still...

One must wonder, dream, and hope.

Foundation Models as digital wilderness

Let us start with these deep learning models trained on large bodies of data. ChatGPT has such a model at its functional core. They have been termed Foundation Models because they “can be adapted to a wide range of downstream tasks.” I have come to think of each such models as repositories of digital wilderness.

What do we do with the wilderness? We explore it, map it, in time settle it and develop it. We cultivate and domesticate it. AI safety researchers call that alignment. The millions of people who have been using ChatGPT are part of that process. We may not think of ourselves in that way, but that, in part, is how OpenAI thinks of us. Even as we pursue our own ends while interacting with ChatGPT, OpenAI is collecting those sessions and will be using them to fine-tune the system, to align it.

Yes, it would be nice to have a system “pre-aligned” before it is released to end users. But I don’t think that’s how things are going to work out. The process by which these engines are trained on large amounts of data is powerful, but it is also messy. While some alignment can be achieved by a small in-house team tweaking the system, ultimately it will have to interact with a much larger body of users. Why not think of alignment as a process concomitant with use?

What are the business implications of a technology where the end users will inevitably playing an important role in fine-tuning and evolving the technology they are using?

Robots and the physical world

One problem that has come up is that of embodiment. These Foundation mMdels, these tracts of digital wilderness, have no access to the real world. Language models are trained on texts, visual models are trained on images. Neither have the capacity to interact with the world.

In an essay he wrote shortly after he took a position as VP of AI for Halodi Robotics, Eric Jang wrote:

Reality has a surprising amount of detail, and I believe that embodied humanoids can be used to index that all that untapped detail into data. Just as web crawlers index the world of bits, humanoid robots will index the world of atoms. If embodiment does end up being a bottleneck for Foundation Models to realize their potential, then humanoid robot companies will stand to win everything.

Those robots will thus be creating more digital wilderness. But also helping to develop it, to align it.

Symbolic Systems and alignment

One issue that has come up is that of the role of the Old School technologies of symbolic systems. Are they obsolete or, on the contrary, will the remain important? While there is wide-spread sentiment that they are obsolete, and that the technology can achieve its fullest flowering simply by scaling up machine learning, I disagree, nor am I alone.

The question is how we are going to integrate symbolic technology with machine learning. One problem is that, traditionally, such systems have been developed “by hand.” The code must be crafted by people who are experts in the application domain. This is costly and time-consuming.

Off-hand I don’t what can be done. Yes, work is being done with hybrid systems where symbolic technology is grafted on to and underlying base of machine learning. I would like to see symbolic capabilities emerge from the underlying artificial neural net technology – perhaps analogous to the way in which language develops in humans, something I’ve discussed in my recent working paper, Relational Nets Over Attractors, A Primer: Part 1, Design for a Mind, Version 2.

I note, though, that I regard the development of symbolic technology as a stream in the overall process of aligning software grounded in vast plots of digital wilderness. It will thus be a gradual process distributed widely thoughout the community of users.

AI as platform, Adept

Two years ago venture capitalist Mark Andreesen talked of AI as a platform, not a feature. He said:

I think that the deeper answer is that there’s an underlying question that I think is an even bigger question about AI that reflects directly on this, which is: Is AI a feature or an architecture? Is AI a feature, we see this with pitches we get now. We get the pitch and it’s like here are the five things my product does, right, points one two three four five and the, oh yeah, number six is AI, right? It’s always number six because it’s the bullet that was added after they created the rest of the deck. Everything is gonna’ kind of have AI sprinkled on it. That’s possible.

We are more believers in a scenario where AI is a platform, an architecture. In the same sense that the mainframe was an architecture or the minicomputer is an architecture, the PC, the internet, the cloud has an architecture. We think AI is the next one of those. And if that’s the case, when there’s an architecture shift in our business, everything above the architecture gets rebuilt from scratch. Because the fundamental assumptions about what you’re building change. You’re no longer building a website or you’re no longer building a mobile app, you’re no longer building any of those things. You’re building an AI engine that is, in the ideal case, just giving you're the answer to whatever the question is. And if that’s the case then basically all applications will change. Along with that all infrastructure will change. Basically the entire industry will turn over again the same way it did with the internet, and the same way it did with mobile and cloud and so if that’s the case then it’s going to be an absolutely explosive....

I agree, and believe that Adept: Useful General Intelligence, is moving in that direction. From their blog:

In practice, we’re building a general system that helps people get things done in front of their computer: a universal collaborator for every knowledge worker. Think of it as an overlay within your computer that works hand-in-hand with you, using the same tools that you do. We all have parts of our job that energize us more than others – with Adept, you’ll be able to focus on the work you most enjoy and ask our model to take on other tasks. For example, you could ask our model to “generate our monthly compliance report” or “draw stairs between these two points in this blueprint” – all using existing software like Airtable, Photoshop, an ATS, Tableau, Twilio to get the job done together. We expect the collaborator to be a good student and highly coachable, becoming more helpful and aligned with every human interaction.

This product vision excites us not only because of how immediately useful it could be to everyone who works in front of a computer, but because we believe this is actually the most practical and safest path to general intelligence. Unlike giant models that generate language or make decisions on their own, ours are much narrower in scope–we’re an interface to existing software tools, making it easier to mitigate issues with bias. And critical to our company is how our product can be a vehicle to learn people’s preferences and integrate human feedback every step of the way.

Judging from what they’ve written, they’re not quite where Andreesen’s conception is, but they’re moving in that direction. All you have to do is put AI at the heart of the application, rather than treating it as an add-on, and you’re there.

Robot Toys, an exercise for the reader

Do some research on the state of robotic toys and companions. Start with, say, the 1990s Tamagotchi, which wasn’t a robot at all, but a little gadget one had to care for. Do a web search on “robot toys.” Think about the Tamagotchi, think about children and dolls/action figures, and think about those robot toys. Take what you find and insert it hear. Perhaps ChatGPT can help you.

Robots and Humans, living and working together

In 1995 Neil Stephenson, published The Diamond Age: Or, A Young Lady's Illustrated Primer. It centers on a young girl, Nell, who is given a precious book, A Young Lady's Illustrated Primer, which she has with her as she grows up. It serves her as both a tutor and a companion.

That’s where we’re headed. But I have no idea when we’ll get there, a century, two, three? Who knows.

At a very young age each child will be given such a book, or perhaps a robot – why not both? – which will stay with them for the rest of their life, functioning variously as a companion, tutor, and workmate, their personal robot companion (PRC).

At the moment time of their third birthday. By this time the child has plenty of experience getting around physically and is getting better with speaking. They will have seen other kids with their PRCs and no doubt have interacted with them. They’ll know that, when the time comes, they’ll be getting one too.

The robot will have to be of an appropriate size, a book too for that matter. I will have to be replaced at the appropriate age. That will require a ritual, as did the original gifting of the PRC, and/or book.

We will evolve toward a society where robots, AIs, and people will be constantly interacting with one another. There where will communities of mixed groups, others of only one kind of being. I have no idea what that will be like. But it does sound like the Fourth Arena will be deeply entrenched in that world.

Beyond AGI, super-intelligence, and the Singularity

What about AGI (artificial general intelligence)? What about it? Originally it was simply AI, and it was right about the corner. But, alas, it really wasn’t. “AGI” was coined in the first decade of the millennium to revivify those old hopes.

While I followed work in AI back in the day, I did it out of a sense of professional obligation. I didn’t think it was going anywhere. The very different AI that we’ve got now IS going somewhere. It’s time to ditch those old dreams in favor, both of current reality and what we can build in the near-term future, and of new dreams.

I feel much the same about the idea of super-intelligence. Oh, I know how the word was used. I could follow the conversations. But I don’t think there is anything there, no substantial conceptual development.

I feel the same about the idea of a Technological Singularity. AGI recursively rewriting its own code until FOOM! super-intelligence? Nah. Not going to happen. Someone termed it The Rapture for Nerds. That sounds about right.

No, something else is afoot. We’re living on the cusp of the Fourth Arena. What could be grander?

Friday, April 29, 2022

Two AI companies converging on (the mythical) AGI from different directions

On Wednesday I blogged about Adept, which is approaching general intelligence by way of creating natural language interfaces for software packages used by end users, in their terms, “a universal collaborator for every knowledge worker.” That makes sense because current systems acquire knowledge, not through hand-coding (like symbolic AI), but through learning. The software environment is ‘native’ for AI engines, whereas the physica world is not, and so they should be able to learn it efficiently so they can act in it.

This morning I found about Halodi Robotics. They’ve just hired Eric Jang as their VP for AI. Here’s what he says:

If your endgame is to build a Foundation Model that train on embodied real-world data, having a real robot that can visit every state and every affordance a human can visit is a tremendous advantage. Halodi has it already, and Tesla is working on theirs. My main priority at Halodi will be initially to train models to solve specific customer problems in mobile manipulation, but also to set the roadmap for AGI: how compressing large amounts of embodied, first-person data from a human-shaped form can give rise to things like general intelligence, theory of mind, and sense of self. [...]

Reality has a surprising amount of detail, and I believe that embodied humanoids can be used to index that all that untapped detail into data. Just as web crawlers index the world of bits, humanoid robots will index the world of atoms.

Jang is right. Reality does have a surprising amount of detail, and an AI engine can’t learn it by reading zillions and jillions of texts. It’s got to get out in the physical world and interact with it.

Here’s what I said two years ago:

So, the AGI of the future, let’s call it GPT-42, will be looking in two directions, toward the world of computers [that’s Adept] and toward the human world [that’s Halodi]. It will be learning in both, but in different styles and to different ends. In its interaction with other artificial computational entities GPT-42 is in its native milieu. In its interaction with us, well, we’ll necessarily be in the driver’s seat.

And, yes, I know, I’ve said that AGI is a chimera, the Philosopher’s Stone of alchemical AI. But if smart and imaginative people take a good run on it, they’re come up with something interesting in the process. Who cares if they don’t make it there. Maybe they’ll make it to Mars instead. They can great Elon when he lands.

Sign me up!

Is AGI currently the Rome toward which all AI research is headed? [The Alchemical Age]

That graphic is from this blog post: All Roads Lead to Rome: The Machine Learning Job Market in 2022, by Eric Jang. From the post:

For instance, Alphabet has so much valuable search engine data capturing human thought and curiosity. Meta records a lot of social intelligence data and personality traits. If they so desired, they could harvest Oculus controller interactions to create trajectories of human behavior, then parlay that knowledge into robotics later on. TikTok has recommendation algorithms that probably understand our subconscious selves better than we understand ourselves. Even random-ass companies like Grammarly and Slack and Riot Games have a unique data moats for human intelligence. Each of these companies could use their business data as a wedge to creating general intelligence, by behavior-cloning human thought and desire itself.

The moat I am personally betting on (by joining Halodi) is a “humanoid robot that is 5 years ahead of what anyone else has”. If your endgame is to build a Foundation Model that train on embodied real-world data, having a real robot that can visit every state and every affordance a human can visit is a tremendous advantage. Halodi has it already, and Tesla is working on theirs. My main priority at Halodi will be initially to train models to solve specific customer problems in mobile manipulation, but also to set the roadmap for AGI: how compressing large amounts of embodied, first-person data from a human-shaped form can give rise to things like general intelligence, theory of mind, and sense of self.

Embodied AI and robotics research has lost some of its luster in recent years, given that large language models can now explain jokes while robots are still doing pick-and-place with unacceptable success rates. But it might be worth taking a contrarian bet that training on the world of bits is not enough, and that Moravec’s Paradox is not a paradox at all, but rather a consequence of us not having solved the “bulk of intelligence”.

Reality has a surprising amount of detail, and I believe that embodied humanoids can be used to index that all that untapped detail into data. Just as web crawlers index the world of bits, humanoid robots will index the world of atoms. If embodiment does end up being a bottleneck for Foundation Models to realize their potential, then humanoid robot companies will stand to win everything.

Yes, reality is (not so) surprisingly detailed, something I talked about in my GPT-3 working paper. I fear that AGI is the Philosopher's Stone of this Alchemical Age of AI that we are still living in.