Essay · Facebook Reality Labs · Codec Avatars
Finding out what people would accept as themselves.
I joined Facebook Reality Labs in Pittsburgh in 2015, on Codec Avatars: photoreal avatars of real people, driven live. The aim at the Lab was to make it “as natural and effortless to interact with people in virtual reality as it is with someone right in front of you”. For six and a half years my job was to find out what the technology could really deliver, and what a person would accept as themselves.
Novel capture and machine learning research pushed the avatars forward, and I understood enough of each step in that system to contribute to it: what the capture could resolve, what the reconstruction did with it, what the solve could then drive. What I brought was a game character creation pipeline, applied to a research problem. Scan data became clean topology, and the facial rig became the instrument I used to author the shapes a likeness needs in motion, handed back to the research team as a target to build toward. That turned what the science could reconstruct into something you could drive, dress and put in front of the person it was made from. Underneath all of it sat one question. What held on a handful of faces we knew, and what would still hold for a stranger.
Quoted from Facebook is building the future of connection with lifelike avatars, Meta, March 2019.
One person, photoreal, and the question is whether they accept it. Avatars 2.0 asks it of billions.
Where it started
Making a person
Scans, doubles, and the jump from a still likeness to one somebody can drive.
Before the capture stages were built we worked with handheld scanners. I learned to run one, to process what came back from it, and to turn that into a digital double that would hold up in an engine. When the facial capture system came online the input got far better, and the problem moved rather than went away. A double is a static thing and a person is not.
So the question moved: how a face built from a scan could be driven by face tracking, in real time, and still be recognisably the person driving it. That was the avatar experience. Everything else was in service of it.
The route was the one a real-time product has to take: a mesh, a set of blendshapes and a rig, driven live in an engine rather than solved offline. A reconstruction that only holds in an offline render tells you what is possible; one that holds live, on a rig somebody can drive, tells you what could ship. Everything I made was a test of the second kind.
It did not stay hypothetical. With a team I helped stand up a working VR experience in Unreal: the avatars in a real scene, driven live from the headset and a mocap suit somebody was wearing. That is a harder test than a render. The avatar has to hold while a person moves inside it, and while someone else in the room is looking at them.
A smile is not one shape. It is one person’s.
Where the gap showed
Making it move like them
Three materials, one demand: motion, look, and the body.
Reality Labs had already published two tests for whether we were getting there: the ego test, whether you accept the avatar as yourself, and the mother test, whether the people closest to you do. They are a good statement of what success feels like. They do not tell you how to find out whether you have it.
So we watched. When a colleague’s avatar moved it was close, and it was not them. Working out why took longer than seeing it. A smile is a mannerism before it is a shape: the particular route one face takes to get there, the asymmetries that belong to it and no other. The expression space we had did not have room for them, so the tracking read the person correctly and had nothing of theirs to drive. The signal was right and it landed on somebody else’s shapes.
Naming it was not the end of the job. “This looks wrong” cannot be built toward, so I went into the digital double’s facial rig and used it to author a more accurate depiction of a few of the people the avatars were built from. Not a description of what was missing: the thing itself, posed by hand, one face at a time. A measurement can tell you a reconstruction is close. Someone has to sit with the face and make the version it should have been.
The look side had the same problem in a different material. I worked with the camera technician on capturing finer detail, particularly skin and the iris. An iris is effectively a fingerprint, no two people carry the same one, and once you have seen a reconstruction wearing a generic one you cannot unsee it.
The same demand ran through the body. Once full-body capture was in place, clothing was the next thing that had to move correctly, and I built cloth simulations to give the researchers something concrete to aim at.
What the research was using
An existing avatar of a lab member, and the shapes it had
What I made alongside it
That person’s own smile, authored on those shapes
Reachable with the shapes we had, or the shapes it would take
What travelled
Making it hold for someone you’ve never met
Findings from people we knew, tested against people we would never meet.
As the research moved toward generating avatars automatically, I took on a different question inside a small art team, alongside an art director and another character artist. Could you drive a stylized version of yourself, or a character that is not human at all?
Stylising someone you know turns out to be hard in one specific way, and it is not a craft problem. Everyone who knew the person loved the result. The person themselves was far more careful about it. That gap is the whole thing. A stylisation that lands for the room can still be a caricature to its subject, and a caricature is a failure even when it is a good likeness.
So I stopped trusting the room. We went back to the people whose avatars we had made and asked them directly, looking for the pattern underneath the individual reactions. Doing it internally meant the loop closed in days rather than months, which is the real luxury of working in a lab.
What came out of it is the part I still use. Recognition by other people and recognition by yourself are two different measurements, and only one of them travels. If a finding rests on people already knowing the subject, it will not survive contact with strangers, and everything we were building was eventually going to meet strangers.
Observed on
Lab members, faces we saw every day
Tested by
Asking the subject, not the room
Holds for a stranger
Needs the viewer to already know them
The move to stylized did not start at Horizon. It started in that lab, on faces that were not photoreal and not human. Later, at Meta Horizon, I answered the same question at full scale: leading hair for a stylized avatar system available to billions of people, none of whom can be asked individually whether they recognise themselves. Knowing which findings travel, and which only work in a room where everyone already knows each other, is what makes that a solvable problem rather than a matter of taste.
Meta Avatars 2.0 · leading hair through launch →