Why Computing with Time Gives Neuromorphic AI an Edge

· EE Times ·

52 min read Original article ↗

In this month’s Brains and Machines, Dr. Ryad Benosman discusses event-based vision, retinal prosthetics, and neuromorphic startups. A co-founder of Prophesee, he explains why he believes AI’s future must come from neuromorphic engineering. Discussion follows with Dr. Giulia D’Angelo from the Czech Technical University in Prague and Professor Ralph Etienne-Cummings of Johns Hopkins University.

[FULL TRANSCRIPT BELOW]

Ryad Benosman
Ryad Benosman

Announcer: Welcome to Brains and Machines, a deep dive into neuromorphic engineering and biologically inspired technology. In this episode, Dr. Ryad Benosman talks about vision theory, computing with time, and neuromorphic startups. Your hosts are Dr. Sunny Bains of University College London and Dr. Giulia D’Angelo of the Czech Technical University in Prague.

Giulia D’Angelo: Welcome to Brains and Machines. I am Giulia D’Angelo…

Sunny Bains: …and I’m Sunny Bains.

D’Angelo: In today’s episode, Sunny is talking about everything from mathematics to prosthetics to the difficulties of getting hardware funded with Dr. Ryad Benosman in New York. After the interview, we will be talking to Ralph Etienne-Cummings from Johns Hopkins University about the issues raised.

Bains: Thanks, Giulia. As you’ll hear, Ryad has been around the neuromorphic scene for almost 20 years. He has an unusual take on the subject, perhaps because he’s a theorist who’s nevertheless worked on problems at all levels of neuromorphic vision—from neuroscience to algorithms, to cameras, to prosthetics.

You’ll hear him talk about his dislike of being put in a disciplinary box, his surprise at becoming an entrepreneur, and his belief that AI’s future is bound to emerge from the neuromorphic approach. There are links to his work and some of the specific papers we’ll be discussing on our website. You can check them out at BrainsandMachines.net.

Ryad Benosman, welcome to Brains and Machines.

Ryad Benosman: Hi, Sunny Bains. Nice to chat with you.

Bains: I’d like you to start today by telling us about your technical background. What did you study, and how did that take you to neuromorphic engineering and vision and all of the stuff you’ve done since?

Benosman: I was interested in how math is built, so I studied the theory of computation, which is considered the basement. The idea is, basically, how do you architect math? How do you ensure math is correct? Do we have hypotheses we’re using and we know they’re not?

And over my master’s, I decided that of the people from my cohort, some of us leave because we think that by going into another field we’ll find new math—we know how to build math, we just need a problem or something to solve. And so after discussing it—I was always a fan of brains, machines, sci-fi, robots—I decided I’d go to do a Ph.D. in robotics, which was quite a cultural shock for me. Being a pure mathematician in that sense, like a pure computational guy, I didn’t know anything about discrete computation, signal processing… and we had basically four months to learn it all and sit for an exam, and the best three got a PhD grant. That’s probably the only time of my life where I studied. I ended up second.

And I ended up in neuromorphic because the only other thing related to the brain was cognitive science, and that didn’t really ring a bell for me. I came in from math. It was too much of a shock. And then vision was hyping, geometry was taking off—this is what gave up to all these special effects and everything you see around. So I thought, “Let me go there and every road leads to Rome, so we’ll get to the brain at some point. Let’s see what they do and how they do it.”

I ended up studying general cameras, which is basically breaking the geometric model they were using—that is, projecting the world on a plane—by getting into more fancy projecting areas like spheres, cylinders. And then breaking the point by making what makes a camera, basically—understanding geometrically why eyes are eyes.

And then, over the years, we could build any eye, and we understood how to build them with mirrors and all kinds of… so, I learned optics along the way, which was fun. And then it occurred to us that the type of images we were getting were not processable with anything we knew. I knew a lot about math, and I knew that this is an unsolved problem. So it occurred, at least to me, that the image was a problem—that’s acquisition. The math was correct, the geometry was correct, the only thing that was not correct was the sensor, basically.

And so I went and started studying the retina from a book called Vision by Rodieck. It’s a wonderful book. If you know nothing, just read it. It’s outdated, but it gives you everything you need to know to challenge the rest, basically. And then I started getting interested in getting to real signals. I discovered it’s called yet another field after optics—robotics. We built robots, by the way. I played RoboCup, I built all sorts of robots, because I like that. But then it passed.

So after optics, robotics, we moved into neuroscience. And I was really keen on getting there and getting the space and starting to contribute. And the only way to do it was to convince people that our hardware—the event sensors that would come later—were the best. And that’s how I got all that effort funded and was able to have a lab that can stretch from recording real signals up to building chips.

Bains: So let’s take a step back. Can you talk about how you got involved with event cameras, and what the problems were with the way they existed at the time and how you processed them?

Benosman: What’s interesting is that when we hit that wall of not being able to process the images we were getting from those real eyes, there was a quest that time was important—asynchronous acquisition was important. So I started working on them from a theoretical point of view, by building a room with 70 cameras that were completely desynchronized. And then studying what time means, what processing with time means, what it means to get delayed information of the same dynamic scene, where is the truth, what it means to add time to vision at the level of the pixel.

And so I bumped into the DVS randomly, because I was spending a week at EPFL and I saw the first working prototype. It was a 64 by 64. And it’s a funny story—the young kid who was doing his master’s on that sensor is now a professor somewhere in Germany. I think I can thank him—sometimes, good luck may change your life or speed it up. It was obvious that it was not the camera I wanted, but one of them.

It was a Friday. On a Tuesday, I ended up in Zurich in Tobi Delbrück’s office. I mean, I knew why this was important for vision—actually, even more than vision, I thought this can apply to anything. I don’t believe there is a dichotomy—brains process the senses the same way. It’s just how you get the type of information. If you look at the periphery of the brain, wherever it’s coming from—whether it’s chemical, photonic—it’s always signaled the same way, with the same dynamics.

And so I started working on those sensors, and they were great. They were not the best, but… and at the time, I remember I was told, “Yeah, people saw these sensors. What you’re saying doesn’t make sense. They’re noisy. They’re useless.” And actually, no, they’re great. And noise is nothing. So at the beginning, it was only me for the first papers. And then I managed to convince people, get funding, get Ph.D. students, and start training people. And we built, I think, what it means to compute with time—which was non-existent. And having a computation and computer vision that had $t$ directly in it, on which you act in matters, is I think our contribution. So we laid the foundation of all this and left empty spaces for people who like to do things their own way.

And then we industrialized the sensor. Because without industrializing it and making it available, no Sony would be around, none of us would be where we are today. So that was an important thing. It’s not fun to do it, but it had to be done.

So I think I’d advocated all these years about how I think these sensors should be processed, and many times I was not happy with what I saw in the field, and people know about my position. When I see events processed as frames, it’s complete nonsense—to this day! I would say it today, I would say it tomorrow.

Bains: So let’s get into that. First of all, we should mention that the company that you founded with co-founders to commercialize the event sensor was Prophesee.

Benosman: Yes.

Bains: But your main contribution to that was not on the hardware side, but in understanding how that information could be understood mathematically—how it could be processed. Can you talk about the difference between processing a spatial signal and a much more temporal signal, as event signals are?

Benosman: One thing: Prophesee was built from a portion of my lab. I think my contribution, in the design, was the introduction of the hybrids. I was the first one to advocate for that in 2008, when Tobi and I were working on this. It was only the three of us, honestly—Raphael Berner, Tobi, and me. Back in the day, I really needed to get some spatial gradient because you can’t integrate after a while, you have a drift. So I said, from time to time, you can have a spatial gradient from something more stable, like a frame. One, it’s good. Two, it’ll give us a second pathway. And three, it will bring everyone to come and play with us. So that was the strategy.

Bains: So what you’re saying is that you are the reason why it was decided to build a sensor that had both event sensing and frame sensing within the same pixels, essentially.

Benosman: Yes. It was built for a European project Tobi and I pushed with a German university for drones. And Tobi wanted to do a color camera. And he was working on event color—basically a differential of rates. That was a beautiful design. So I told him, what if we put APS pixels that would give me gradients? I was computing optical flow back in the time, and it was not very stable—the project didn’t make it. But Rafa Berner really listened, and he did it, and it’s in his PhD; that’s his contribution. There’s lots of problems in that design, in the sense of combining both. And I think over the years people understood them.

So I think hardware is great. I love hardware. But hardware makes sense if you know what you need it for. If you build something regardless of what it’s used for, you can build the most beautiful thing, but it’s still useless. It has to be computationally efficient.

So I think I see the entire stack. I just don’t touch the analog side. It’s too complicated. It’s like Roger Federer, a professional tennis player—you have to do it in the morning and sleep and still do it the day after. I don’t have the skills or time for that.

Bains: So talk about this side that you are best known for, which is the mathematical side of that—and how it looks to be processing in time rather than processing in space, at least most of the time.

Benosman: Imagine the world as it is. There is no time, right? Information flows all the time. It’s changing, it’s continuous. And there’s a beautiful paper called the Plenoptic Function. It’s a very old paper. It was one of my favorites during my PhD. It just shows, basically, that the world is a set of infinite viewpoints where you can grasp light from every direction, infinitely. I think that model is correct. It’s beautiful. It’s called the plenoptic function, it has 12 parameters, and I think that’s the mother of all cameras. Whatever you do, you’re evolving in that space. Infinite points, infinite directions of light from every point in direction.

So now if you stop at one point, what do you need? How do you get information through one ray of light? When we started computing, people thought, “Yeah, we have frames. That’s what brains do. I see frames.” No. There is no frame in any brain. Nobody has ever seen an image anywhere. So we don’t know—maybe frames don’t even exist. We know there is integrated information, but it doesn’t mean there is a frame. So that’s the main challenge.

And the DVS is one of the methodologies; you can decide to take that continuous signal of light in one direction and chop it, saying, “Okay, I have silicon, my dynamic range is limited. How can I fit all that light inside with the best compromise? I’m going to do it as a differential.” That was a super smart design, to make that circuit stable and with a good yield. It’s really nice. It’s beautiful.

But, again, think about it—you have infinite ways of acquiring information. Brains have probably 15 features out there, or 50, or other animals. Everyone is getting information from that stream of light according to his needs. So I cannot tell you there’s one function, there’s one DVS, there’s the ultimate camera. That function changes.

So that’s how I took the signal. I thought, “Okay, they chose that. This is not what I want, but I’ve got time. That’s all I want.” And so, based on time, you get lots of information. Would you like to see the world at 20 megapixels at one frame per second? Or would you like to see the world through a doorknob at 10,000 frames per second? I would take the 10,000 frames per second—I’d even go and play the French Open. I’m going to beat everyone because I’m ten times a better predicting machine. So that’s how I saw it.

And so all the compute came from the fact that, at the end of the day, when you think of whatever type of information, it’s going to come as sorts of events. I like this metaphor of a castle. You have a castle, you have to defend it, so you put soldiers around the walls. If you were a conventional system, you would put the drummer, and every drum beat—every soldier, starting from their number you assign to them—would send you his information of exactly what he sees. So he would tell you, “I see a mountain. I see a forest. I see a mountain, a forest, a mountain.” The same thing again and again. Well, now you’re being attacked. And so, “I see a mountain, I see a knight!” But you have to hear through all the unnecessary things that are happening. “I see a knight, I see a mountain… oh, I think there are two of them!” What!?

So, at the end of the day, you’ll be attacked. You’re slow, you’re inefficient. It’s noisy. The place is unlivable. So what would you do? It’s common sense. You go up there and you say, “Guys, I’m interested in horses, weaponry. That’s all I want to hear.” And then everybody goes home and it’s quiet. And when somebody fires—10—if he screams, I can probably ask if 10 is reacting to horses. 11 is reacting to a sword, 12 is reacting to a catapult—according to who fires, I know where it’s coming from. He doesn’t need to explain anything to me. I got it. That’s what brains do.

And so you save lots of information, you have dynamics, and you have the same type of information that you have in a frame. More crude—you can integrate it if you’d like—but you have its dynamic. You see how it’s unfolding. So the features are not spatial. They are spatial and temporal.

Bains: That’s a nice metaphorical explanation, but can you explain a little bit more about time surfaces and how you handle them in your work? Because that’s a key concept for you.

Benosman: Time surfaces is this idea of seeing events as music. You write music in the score, and people have discretized time—we know we have a tempo, it tells you of the time resolution. And in the middle of the time resolution, you have a half tone or whatever. We fix the time. It’s discretized, let’s say it that way.

Whereas event music is if you take eight pixels and you put A, B, C, D notes on every note, and you let the music go, you have music. Every character has some music to it. Everything is encoded—not only about what shape is passing by, but also at what speed. It’s music.

So time surfaces—which is a joint effort with Bert Shi, by the way; it was a time when I was in Hong Kong and I was showing up at his office every other day—it was a way to basically contract it. Contract the writing of a continuous time representation into something we can play with like a vector.

But the paper, really, is not about time surfaces. In the paper, HOTS, time surfaces is a simplification of a more complicated computation you can find in STICK—which is another paper. But if you look, what really matters is that HOTS is showing, basically: here’s a way of expressing dynamics in a vector on 5 milliseconds, and I’m going to find you the most representative one. And I’m behind, and I’m watching all these notes. So if somebody’s playing one type of note, and we’re learning how those notes arrange together into, let’s say, some nice tune of 5 milliseconds—so you have lots of small tunes playing around. And then you have something else behind that takes all those sets of tunes that you’ve decided are important—how they align over time, also with delays, also with arrangement—so it’s a combination of a combination. And the deeper you go, a feature tells you this is a combination of a combination. It’s integrating a larger timescale. That’s the beauty of that network. Time surfaces are what makes it possible, but instead of taking a box of events and saying, “How am I going to express this and show to the world what’s a feature inside the box?” I think the way we’ve done it is more inspired by braids. It chopped it into different scales of dynamic and linked them together in a very compact way.

And we used HOTS all over the place, even on real neural signals. It works extremely well in tracing neurons and finding connections and understanding receptive fields. It’s a way of expressing anything that unfolds over time—a continuous representation of music.

I love music. I’m a cellist. I’ve played music all my life, so probably I’m biased. But honestly, with event sensors, if you really want to play and do it well, it’s music. It’s timing, it’s information that comes—simple. And so you have everything people do in computer vision, except this time you add one more dimension, which is time. So instead of getting an edge, you get a beautiful space-time edge moving. You can get the edge, you can get what type of edge it is, how fast it’s going, get more information about it.

So that’s, in my opinion, what we introduced. And for the first time, when you think that way, you can start writing papers with $t$ on them; time. Because that’s what matters. And not even time—the time between events. That’s where information is. And then how you handle it, with the specific scope of never, ever smashing it into a frame! Because you lose everything.

Bains: Absolutely. So, a big part of your work has involved thinking about how to create a visual prosthetic that allows the users to see with whatever part of the visual system is physically intact. For instance, you use the retinal ganglion cells to convey information to the brain, even when the photo detectors—so the rods and cones—have died. Can you explain how that works, first of all?

Benosman: Yes. So the retina is a seven-layered neuron circuit. It’s the only approachable part of the brain. You can take it out, put it in a dish, shine light on it, get every signal firing of every neural cell. And although I’ve been out of academia for three years, to my knowledge, there is no model that fits the data. For a long time, it was considered as a simple filter. It’s not. It’s just doing a tremendous job. It’s just genetically wired. There’s no learning, there’s no feedback, at least not for humans.

So the retina is an inverted sensor. The photoreceptors are underneath. The entire circuitry is up. And the further away you get from the photoreceptors, the more the information is processed, features are being created over space and time. And more complicated features—there’s inhibition at the level of layers. It’s a very complex machinery. And then what you get out of it at the optic nerve, where every ganglion cell—we think there are 15, some say 25, we don’t really know—but you have 25 time surfaces, or descriptions of a space-time activity, happening in some receptive field. So some of them react to moving edges, some react to zoom in, zoom out, lighting. All sorts of incredible retinas exist in nature. There’s not a single one that looks like the other one.

Bains: But you also need to get to how the optogenetic part works.

Benosman: Yeah. When the photoreceptors are gone, you have a circuit. So they don’t die, all of them, at the same time. It’s a long process. So what happens is you find yourself with a working circuit, but there is no more photonic transduction. So the scope is—by electricity or optogenetics or different ways of making the neuron fire again, the idea is—what if I can substitute myself to the photoreceptor that just died and excite the remaining circuit? So obviously, the closer you are to where the photoreceptor was, the easier it’s going to be to inject the message in there. And the further you are, if you go to the ganglion cell that formed the optic nerve, the messaging and the compute is tremendous already. So where do you want to inject the message? That’s the real game. And I’ve injected both up and down. I’ve worked epiretinal, subretinal, and also did optogenetics—three devices we worked on.

Bains: And the optogenetic work, as I understand, involves using a virus to make the RGCs emit rhodopsin, which makes them light-sensitive. Is that right?

Benosman: Yes. It’s a gene from an alga people isolate. And then the game is to find a good virus that you naturalize, remove the bad code, and you can target specific cell types. That’s the entire game. Then, once the virus delivers the message, what message did it deliver? That’s another game. There are lots of messages. DNA, you can put inside cells. So you must know that this is irreversible—meaning that if you got the virus, you cannot get another one after. But that’s the main idea. So there are lots of people working on how to deliver the message, and what’s the best message.

For now, these opsins are, on paper, really great. But they’re slow and they need lots of light. The best one we could use was supposedly the lowest photonic requirement. But when you push yourself outside the eyeball and you start sending light to send that active amount of photons to trigger the neuron to fire, we had to basically solve two challenges: How do we encode information? What does it mean to send information to a neuron through light? And second, how on earth are we going to send the amount of light into that eyeball and make it fire?

Bains: Did you say it was about 10 times the amount of light? Or was it more?

Benosman: Seven times the power of light in the desert at noon.

Bains: Wow. And how does that compare with the amount of light that your rods and cones would—is it three orders of magnitude?

Benosman: Seven.

Bains: Seven orders of magnitude! My goodness.

Benosman: If you wear these glasses that we’ve built and you’re healthy, you just lose your vision instantly. It is gone.

Bains: Right.

Benosman: These are strong lasers we have used. But the endeavor is not that. It’s to make it a medically approved device.

Bains: Yeah.

Benosman: It’s approved in the U.S. and Europe. That was the challenge. And the only way to do it was with event sensors. And the only way to put information in that circuit was to put some temporal dynamics—otherwise people would see flashes. So the fact that we were able to transduce information into time made it really workable.

Bains: So by working with the event sensors, you’re able to massively reduce the amount of light that you need to communicate the information that you need to communicate.

Benosman: If you look at the FDA or European requirements, you have a certain amount of photons that you can send inside the eyeball—then, it has never been done. We had to find an agreement on what it meant—was it toxic? So we did it twice with two implants inside the human eye, like electricity, and using the same system. We also have a projector that—instead of optogenetics, we did the PRIMA device, which is photodiodes. So it’s like a solar panel where you shine light on the solar panel and it gets electricity out and stimulates cells. It’s roughly the same projector of the optogenetics, although the requirements were different and, actually, it required less light than the optogenetics.

It was a terrific endeavor. It paid for all our neuroscience questions and more. And after, as I said, doing optics and understanding—we even modernized those opsins. You can find papers where we understand their dynamics, because if you make a neuron fire more than three times, it just passes out. You can’t do that.

Bains: So your contribution was all about really that modeling side of understanding what the visual system was going to do, understanding what the opsins were going to do, so that you knew exactly what code to send into the eye through whatever mechanism to communicate the information you wanted to communicate.

Benosman: Yeah. What set of strategies we could think we could use. And you can make tests where you can see—for example, on an implant—if it saccades in the right direction if you hit the implant. But primates don’t talk, so you have to go on a human…

For the microchip, we had modes [established procedures]. Everybody has modes, one, and two, we also built the entire stack—going from the goggles to the code to building the safety devices. It was really high-level optics, sending this seven times the power of the sun. We had to reinvent tests and so on. Half of my effort—my personal effort for the last decade, probably—was on building systems like this. Which is fine, because you get to see the real signals. But it was pioneering work.

Bains: I understand there was a clinical trial related to this.

Benosman: All of them had clinical trials. Yes.

Bains: Can you talk about the results of those?

Benosman: In general, I think there’s a polemic around this work, especially the optogenetic one. I won’t get into that. But I can say that after really building devices, after doing epiretinal, I’ve done subretinal, we did optogenetics, we also studied on how to do it in V1. We were in the DARPA—I don’t remember the name—to build one for the V1, and we were shut down after a year.

I think there is so much subtlety in how two neurons talk to each other, and electrodes are 100 microns by 100 microns. You already have hundreds of neurons. As you see in one of our papers, if you just fire it like that, especially in the visual cortex where columns are highly specialized, it just cancels everything out. It’s not going to work. So it can work on the retina where the columns are very clean, they do the same thing, there’s not lots of diversity. It’s like going in the middle of Central Park and putting a huge set of microphones and just sending 2,000 decibels in the park. Everybody is going to scream. At the end of town, they’re going to hear them. But it doesn’t mean they said something valuable. It doesn’t mean anything. You’re just hitting everybody in that column.

So it works. It depends, also, on how the implant is. It depends on how cognitive the person is—you have to relearn to learn, it’s not the vision. It’s not going to work right away. As long as I don’t see a technology that allows me to talk to every synapse—which, in that case, I would be happy to return and finish the job—I think it’s useless to inject the message in a brain today with what we have. However, you can decode easier. And that’s what I did in Pittsburgh with Andy Schwartz, where we decoded M1 and worked on lots of decoding or learning. We had enough methodology to bring to neuroscience to tackle the same data in a different way. So if you ask me why I did that, I think it’s a joint effort with neuromorphic—it goes together. If we understand brains, we understand how to talk to brains, then anything we get about the brain will make us better in being neuromorphic.

Everybody thinks about the brain. We had the unique—I mean, this is luck, right? We did two clinical trials on two retinal implants, optogenetics, we did all of it. And, at the end of the day, you realize that you had the opportunity at least to try to talk with a brain. Just a bit. And I don’t think it will happen, at least not in my lifetime, unless we see a big shift in materials. It’s too far.

Bains: You’ve worked with a lot of different companies in your career, including co-founding Prophesee, your work with GenSight—who you were doing the prosthetics work with—and even Meta. I know you can’t talk about the specifics, probably, about any of that work, but can you say what you’ve learned from these industrial commercial experiences?

Benosman: Believe it or not, I never thought I would have a… I don’t know if you call it a career in entrepreneurship. But I discovered, after doing it, that I know how to do this. It was never my scope, my aim. The first time I put money in a company, it was random. Then, Pixium was based on our work, and they asked us to put in some money, and I discovered what it is to be in a startup. And then we did the GeneSight work, all of it, all the component coding, all—it was a huge team; we had around 45 to 50 people in Paris, it’s a huge lab. We had over millions we couldn’t even spend, because the university didn’t let them spend. Anyway, it was a mess! We were not at the right place. And it’s funny because you think, “Okay, I’m not going to do it again.” And then we did GeneSight, and I worked on Pixium and all that, because I wanted to get event cameras to be industrialized.

To create Prophesee, we needed the camera I thought was the best. And there was Christoph Posch and Daniel Matolin in Austria, and they were at the end of their time in that institute, so I raised money to fund their place to bring them. Companies allow you to do stuff like that. And this is how I got the hardware team in my lab and then spun it out. But if I didn’t convince them that we need that sensor to do the retina implant, we would never have Prophesee today.

So I think companies are great for lots of things—executing on an idea, if it’s very clear. Because people have to understand, everybody thinks he’s going to end up being Elon Musk, but the reality is a startup is losing money every day. So the freedom you have to really change the world is very thin. The game has changed recently, I can tell you about it, but back in the day, if I stayed in my lab thinking about event cameras, I’d probably retire and it would still be in papers. The fact that we took that sensor, industrialized it—go for it! Like with the implant, we had no idea what we were doing. I mean, sealing a piece of silicon to put it in a human being… I pray for those who do it today. You have no idea what you have to go through! Even sometimes you have to go to the same foundry five times to get the wire put correctly, or you’re going to stick it into a human being. You have to go on pigs and you have to go on primates, and sometimes you can switch on the other. And you don’t want to do that. And you have to sometimes even create the tools to put the implant inside—which we had to do. Companies allow that. If I was thinking about implants in my life, we’ve done three campaigns. Starting from scratch, I would never have done it without $30 million or $100 million, whichever can happen.

So that’s how companies, at least in my opinion, should be seen by scientists. It’s a way for you to see that there is life after a paper. A paper is good, but a paper should work seven days a week, 24 hours a day. And a paper cannot do that. And you, as a scientist, will never have the manpower—it’s impossible to scale it. You can, but it’s going to take you forever. So at some point you have to industrialize it, let it go, and make it available. Somehow, we created Prophesee, and now event cameras are everywhere. You can buy them for $300 today. It’s good. It’s fun.

Then you have the startups where you have to build hardware. Those are very hard to fund because nobody wants to fund hardware. Finding a way to create a Prophesee is a miracle. So creating a Brainiac that ended up being GrAI Matter Labs is even more of a miracle. It happened because there’s a guy at DARPA who read one of our papers who thought there was a machine hidden behind, and he asked us to build it. And we became good friends, by the way. His name is Phillip Alvelda; if you listen, he’s a good friend now. So, it’s really a quest.

GrAI Matter Labs is a hardware company. You can’t fund it in the U.S. We managed to fund it in France. That’s—wow. And sold it, even sold it!

So, anyway, have one idea pushed and industrialized when you think it’s needed. I never started any startup to make money. I never thought I would make money out of startups, nor even sell one. I thought if I can get those cameras to work and get better and better processors and have somebody building them for us and not me with my own lab, that’s going to be a big win. So when you see them built by Sony and licensed, you say, “Yeah, we did a good job. Probably you can do better.”

And then, people have to understand that the companies are not yours. After a while, people don’t even know you funded them—sometimes they even remove you from the founders list because you’re not happy with what they do. So it’s a very love-hate relationship. It’s going to be beautiful at the beginning, but I doubt a scientist should be happy after two or three years, because then it’s too far, and sometimes people are too far away from the original idea. But they will still get somewhere. And keep in mind that you have a board and you have to convince the board that what you’re doing is right, and you’re praying for the next cash because you know that every day you’re losing… so it’s a very uncomfortable situation. I say to anyone, if you are not ready to wake up in the middle of the night and sweat thinking, “How am I going to pay everyone tomorrow, including myself?” then don’t do it. It’s horrifying.

So then I moved to Meta. It was for family reasons—I needed, really, to be in New York, and I think my time in Pittsburgh was over for different reasons. I thought, you know what, I’d love to see how those companies work, what it means to be in heaven. Because everyone I knew told me it’s heaven. Lots of my friends are working at Meta, or Google, or… So I asked, and I ended up at Meta. I have lots of friends there, and I was really staggered by what I saw. I’ve never seen something like that. As an academic, it’s quite interesting to see, because if some of you go there and show your last paper, so proud, showing 99.9%, they are not interested in one nine. They want five nines. So the level of engineering they do is far beyond what we can think of—in means and quality of people. Everybody there is an A-plus. And technical! They hire extremely talented people. They have extremely smart guys leading the teams. And it’s a different way of doing research.

I was lucky enough to see all of it. I was high enough. I really saw how it works. I’ve seen things that sometimes I didn’t know were even feasible, to be honest with you—or precision I didn’t think could be attainable by human beings, stuff like this. So it gave me a perspective about how they work, and the means they have, the money, the people, how they’re organized. It’s really interesting. I liked it. I left it for different reasons, but I would’ve loved to stay. I still have good friends—some people there I’ve done my PhD with in machine learning. I’m from the lab where Yann LeCun and all that line came from, although I was not doing machine learning there.

And so I wanted—after eight startups, one GAFA company—I wanted the challenge of building a company where I don’t rely on VCs. And that’s why I left Meta. Because I saw how I could do it.

Bains: So finally, you’ve been heavily involved with the neuromorphic engineering community ever since you started with it. What was that? In the early 2000s?

Benosman: 2007 I showed up, I think, in Tobi’s office.

Bains: Okay. So can you tell me what it’s been like to be involved with neuromorphic people, and how you see the future of the field?

Benosman: Honestly, when I went to the INI [Institute of Neuroinformatics] the first time, I was mesmerized. I’ve never seen anything like that. I’ve never seen people of that quality all at the same place at that level. Rodney Douglas is an amazing guy—he’s opinionated, you can say, whatever, but we owe him a lot. Just seeing that showed me what I should build, basically.

I was in my mid-30s, and it was the right time to see that something else exists—the couches, the cool side. It’s cool to be a neuromorph, right? And they’re asking real questions. If you go to computer vision, it’s mainly engineering problems—it’s not really interesting, in the sense that it’s fun. If you like crosswords, you’re a specialist of crosswords. If you’re not a crossword expert, you probably want literature, because the words originate from there. It’s larger, it’s science. You touch the science, you touch the unknown, you try to explain nature—with hardware and math and brains and flesh and crazy people.

And it was fun to be a neuromorph back then, I thought. The first time I went to Telluride, it was a shock. People in shorts—like, you used to see people in conferences with a tie. Take Terry Sejnowski, for example; he showed up and you can chat with him around a beer. Everybody is accessible, ideas flow. Nobody shouts because they don’t agree with you. No problem is declared solved by anyone. Actually, nobody agreed on anything—to the extent that one summer, after the word “cognition” has been abused and used, we decided to find a common definition of it. And this is fun! This is science. This is what makes science fun. It’s not about cameras. It’s about real ideas, like: how do brains work? That’s, I think, what I like there in every way possible.

It’s basically an anti-disciplinary place, which I love. I never understood disciplines. In Paris, my lab could do optics, we had an optic authorization, lasers, we could design circuits, we could record from brains—so we were like 10 disciplines in one, and nobody wanted us because nobody thought we belonged there. It’s stupid. Neuromorphic is what science should be, basically. Open to engineering, who are open to understand that engineering is good, but the real signal is even better—it would teach you to do better engineering. Take the machine learning effort. 99% of them, probably 99.9% of them, have never seen a real neuron firing a real signal. What is noise? Why are the neurons inside M1 so huge and firing all the time? And people think it’s rate-coded, and it’s not actually—it’s temporal coding in there. It’s smart, it’s beautiful.

So I think it brings engineering to become a science. This is what I like. And for somebody like me—who likes to build things, innovate—you have everything. You can build stuff, you can look at this stuff, you can talk to crazy guys. When I was in the Human Brain Project, there was a pillar that talked about consciousness. It’s not consciousness, I think, where it works like this. No. It’s based on data, on people who went into a coma. People trying to measure something—measure—not explain. And this is the beginning of science. If you cannot measure, what am I talking about?

I don’t know, maybe because I’m not an engineer and I come from hard sciences and I always had open questions, I felt at home. I’ve never felt at home in computer vision. I’ve never felt at home in robotics. But neuromorphic engineering—if you want to be at the cross of all this—is home. It’s the place to be.

Bains: And what about the future?

Benosman: Listen, you asked me about startups. I’m going to link it to startups. Years ago, you raised $1 million, $2 million, you can do Prophesee. And then if you sold it for $300 million, you scream and say, “Hi, well done!” Now, people have to understand that VCs are… it’s a casino game. You take $3 and you have to give them back, pay yourself, and probably more. But the minimum is to bring back three times the dollar you got. So the game has shifted. Now, with AI, companies require more money, and at the blackjack table, now, tickets are $500 million, $1 billion. Companies like Prophesee are harder to fund—not that model. Now, people want $1 trillion valuation, $3 billion valuation, which is crazy.

So, why are they doing this? Because they’re looking for an alternative for AI. And AI hit a wall a long time ago. Everybody knows it, the founders know it, all of us know it. But it’s working so well—although imperfect—so well already that it gives the illusion it’s doing the job. But it’s not. It’s not taking us anywhere. It’s not only me saying this—all of us say this.

And so neuromorphic is, for me, the only place it’s going to come from. These are the only guys in the world who think about something that looks like the real thing. So unless there’s somebody who is hit by lightning tomorrow morning, who comes up with a magnificent idea—I don’t think it exists. Because you need to create new math. It’s not going to come from nowhere. We need something new, and we need something that is asynchronous. And we need something light and low power, and we need something that thinks about the fact that computation is not infinite, and so it’s tied to power budget. And we need somebody who understands what it means to communicate in saltwater within the time domain. That’s why we create music, and blah, and blah.

So the odds of something like that coming from another place is, for me, extremely low. It can only come from that. Or maybe from a neuroscientist who is handy enough and sees something we don’t see. It cannot come from what we know. What we know can only produce what we know. It’s useless.

Bains: Ryad Benosman, thank you so much for coming onto Brains and Machines.

Benosman: Thank you.

D’Angelo: Thank you, Sunny. I love that he was open to sharing the struggle of running a startup, and that he gave us an overview of his incredibly long and successful career. For more about Ryad Benosman’s work, please go to BrainsandMachines.net.

And now we welcome back our regular commentator, Prof. Ralph Etienne-Cummings from Johns Hopkins University.

Ralph Etienne-Cummings: Hi guys. How are you doing? Happy Spring!

Bains: Finally!

D’Angelo: Ciao, guys. So, let’s start with Ralph’s impressions, and then we will go ahead.

Etienne-Cummings: I loved the interview. I like the fact that he talked about the mathematics, the hardware, the application to real systems—such as the retinal implants and retinal prosthesis that he was doing. But my favorite thing he said, which I thought is really true, is the fact that event signal processing is applicable to everything—not just vision. I think that’s a really key thing. And now we’re starting to see that protrusion, if you will, into tactile, into audition, into other kinds of things. So that’s a really important aspect of the work that he does.

D’Angelo: Yeah. For me, it was super enjoyable to listen to. One thing that is super important for me, and I’m happy that he said it: “The retina is genetically wired. There is no learning there, and there is no feedback.” And I thank him so much for saying this—of course, it’s not the same learning that we think for complex scenarios. The beauty of what he said is also the importance of biology; that we need to understand biology—and we need to understand the brain—to then talk to the brain, which is super important.

Also, one other thing that I think we don’t think about enough as researchers is the fact that there is life after a paper. We usually just think about getting out with the new paper, but we should commercialize a little bit more of what we do. I also had no idea that the hybrid sensor was coming from him. It makes total sense, and it’s really nice to know.

Do you have any impressions that you want to share, Sunny?

Bains: I do have one thing I wanted to say, which is that when I was preparing for the interview, I looked at a lot of his papers, and they are very mathematical. But one of the things that was nice about that—as I’m discovering, very late in life, a kind of love of theory—is that when you analyze mathematically, you can see where the trade-offs are. So you can see where he’s trading off density versus sparsity, where he’s trading off using a fixed timescale versus using a flexible timescale in order to make a problem tractable. And that, as I read these papers, I kept thinking, “If we had more of this, we would understand all these engineering problems a lot better.” It was just really nice to be getting a sense of where these trade-offs are.

Etienne-Cummings: And for me, trade-offs, mathematics, and the theory are all absolutely crucial. I think that’s one of the limiting factors of neuromorphic engineering, if you will. We are a little bit too phenomenological. We see something and we’re like, “Oh, let’s see whether that does something.” But to have a mathematical underpinning of why, and what it means, and what are the bounds, and what are the capabilities that come with it—I think is really crucial. That’s why seeing some of the work that some of the control theorists bring to this domain is really important. Because they always think in terms of, “What are the bounds of functionality? What is optimal? How do you reach those bounds?” and so on. I think that’s the part of neuromorphic that we’re still missing a little bit. So I think the math is super crucial to make that happen.

In terms of the specifics of the items that he talked about—like the prosthesis talking to the cells—I think that is so crucial to understand that the retina is already wired. So if you are going to tap into the retina, how do you make sure that the signal that you’re sending up the pipe, so to speak, is the right signal? If I’m taking the ganglion cell and I’m hammering it with a signal to say ‘fire,’ what does that mean by the time it reaches the visual cortex?

So those kinds of things are really difficult to understand. And that’s maybe some of the reasons why some of our efforts in prosthesis have not been as successful as they could be.

D’Angelo: Yes. On this, I have one big point. But first, I’d like to mention “HOTS: A Hierarchy of Event-Based Time-Surfaces for Pattern Recognition,” which is an important, beautiful paper that comes from Ryad. And then I need to tell you this anecdote: I came here to the Czech Republic, and there is a beautiful seminar here from Ján Antolík—so I joined the seminar, because it was talking about the retina, and then the V1 and the cortical mapping, and I read “Ryad Benosman: How to talk the same language as the brain.” And I was so happy! Ján Antolík is a professor here at Charles University in Prague, and he collaborated with him exactly on what you were just mentioning before, on the V1 prosthetic vision restoration using a large-scale spiking model of a cat, actually—V1 cat—with more than 100K neurons mimicking the V1 endogenous feedback. And actually what is very beautiful is that this validates what Ryad was saying—that the cortical stimulation needs brain-like temporal dynamics. So this is really important for me, because this is exactly where I would like to go next with my research—a secret that is not a secret anymore. But I think that this should be really the ‘what’s next’ with this event-based and SNN computation, specifically about rating algorithms. What do you think?

Etienne-Cummings: I’m gonna take a slightly different perspective…

D’Angelo: I knew it!

Etienne-Cummings: …and it has to do with the fact that I think, if we’re talking about retinal prosthesis in particular—not general stimulation or talking to the brain, but retinal prosthesis—I think the next step is actually in the organoid space. Which is: you take a piece of your own DNA and you can then generate a piece of your own retina. You can basically then cut out the parts of the retina that aren’t working and implant those directly into that. Now, why do you do that? You do that because, firstly, you’re not going to have all the rejections that are associated with either putting silicon back there or doing some kind of virus manipulation with optogenetics. Secondly is the fact you can actually restore some of these photoreceptors that might die—that are being killed. But here is where I think the event-based part still comes into play: when these retinal organoids are produced, there are no connections. Why? Because they didn’t have all the activity that happens pre-birth; the waves of signals that propagate through a fetus’ retina that forces connectivity and gets the right connection between different layers of the retina. Here is where we could do the type of processing that you’re talking about, Giulia, which is connecting stimulators and recorders directly to the development of these organoid retinas and then get the piece that is accurate to be put into the retina.

Bains: That’s really interesting, Ralph. I don’t think I knew that retinas were developed epigenetically.

Etienne-Cummings: Yeah, totally. There’s a guy here at Johns Hopkins, I forget his name now, who’s very famous, but he’s gotten cells to fire and structures to form. There’s a woman that I used to work with before she moved to the University of Denver; her name is Valeria Canto-Soler, who has done some of the earlier work in this domain as well. It’s really interesting stuff.

D’Angelo: But then let me poke you. First, I’d like to read the title of the paper, because I think it’s really important: “Assessment of optogenetically-driven strategies for prosthetic restoration of cortical vision in large-scale neural simulation of V1.” So this is very good, and we will put it into the notes of the podcast. Let me poke you. Is that easy to control? Is that easy to manipulate? Because, for example, one of my latest papers is around amacrine cells, which for me was so easy because it’s just basically based on two Gaussians and then the subtraction of it. Can you really do something that easy—that was easy for me—in what you were explaining before?

Etienne-Cummings: It’s still under development, right? And what I mean by that is this is still ongoing research. So you get the growth of bipolar cells, you get the growth of the horizontal cells. The subtraction, as you indicate—which is basically the smooth Gauss versus a sharp Gauss—that is not, at least to my knowledge, I haven’t seen that being completely reproduced in the organoid retinas. But, look, that’s exactly the kind of thing that they’re trying to get at. And then the temporal responses, or the temporal dynamics that come in the inner plexiform layer, and that then goes to how the ganglion cells modulate. So there’s a lot of actual structure.

So yes, it’s predetermined, but there’s a lot of activity-driven things that are also part of development. And then ultimately forming that connection is going to be key. And if you can do that, I posit that then the retinal disease mitigation gets a little bit easier—but maybe not completely solved.

D’Angelo: I see. There’s another interesting point to add here. For example, there has been found a link to the shrinking and enlarging of the horizontal cells—and this changes, of course, the incoming light and then the perception of the visual information. And there is a strong link with the fact that there are no schizophrenic patients that have this problem. So schizophrenia could be linked to a vision problem. This is very interesting. There is a recent 2025 Nature paper that talks about this, and this would be very beautiful to see if any improvement can be done in this regard.

But anyway, on another note, I just have a curiosity: why does nobody want to fund hardware?

Etienne-Cummings: Because hardware is super expensive. Sunny and I were talking about it just the other day, in fact. The cost of getting the state-of-the-art processors is approaching tens to hundreds of millions of dollars to start. That’s the non-recurring engineering cost. If you want to go back to the old processors, like 180 or 65 nm, then it gets a little bit more affordable. But then you cannot put as much processing in the same area. So it’s just the fact that it’s a really expensive endeavor. There’s a lot of error that comes about. And in fact, I think Ryad talked about that, right? There are multiple cycles before you get the thing to actually work. So each time, you are burning funds.

And the other part is the actual human cost. It takes a lot of time and a lot of expertise—and these are expensive engineers—to be able to do it all. So when you put it all together, it’s much easier to write code than to do hardware!

Bains: Absolutely. And I wanted to also say something about this issue of going smaller and smaller and smaller. Now, we as physicists and electrical engineering people have been told that Moore’s Law was gonna run out for decades. But the thing is that even though we are making smaller and smaller feature sizes, the fabs are also becoming almost exponentially more expensive. So my question really for the industry is: why are we focusing on smaller and not focusing more on diversifying?

For instance, I was reading about some of the work that Charlotte Frenkel had done with FDSOI—with silicon-on-insulator technology—which solves some of the power problems, because it solves the leakage problem for these very small circuits. So I’m wondering why we couldn’t just think about getting more creative and doing things at slightly bigger nodes, maybe, than we could—but doing them in more interesting ways that would support a bigger diversity of things, that would work better for a bigger diversity of problems.

Etienne-Cummings: Totally. I think as academics—which is the angle that I’m coming from—that is the only way we can stay in the game. By basically saying, look, you’re going to use older processors, bigger processors that are allowing us to then maybe try out new algorithms, new architectures, new ideas. And then of course maybe we can connect them—or hybridize them with optics, hybridize them with other capabilities that can make it a little bit newer in that aspect. That you’re going in some other direction, as opposed to just pure CMOS silicon down.

But from a commercial perspective, the place where you’d hopefully get the most bang for your buck when you sell these products is if I can put as much computation in a small area as possible. So that naturally forces us towards smaller and newer technologies.

But there is that trade-off that you’re talking about, which is: nothing is for free, either on the manufacturing side or on the use side—which is the power and the leakage and all these things that you’re referring to, Sunny—all of that comes at a cost.

Bains: But all of that edge stuff… if we think about all of the devices that could be smarter—smarter so that they use less energy, smarter so that they do things in a more efficient way—all of that intelligence, it doesn’t all need to be teeny tiny. It’s not all going into an earbud or into a mobile phone or something like that.

One other thing I wanted to say on hardware: we were comparing the cost of process nodes. So, very specifically, we were talking about the fact that if you’re working on a 28-nm process node, you would be expecting to spend on the order of $3 million, right?

Etienne-Cummings: As non-engineering?

Bains: Yeah, that’s for masks and for the design cost. Whereas if you’re working at 180 nm, you’re looking at hundreds of thousands.

Etienne-Cummings: Tens of thousands to $100k.

Bains: Really?

Etienne-Cummings: Yeah, it’s much less. You can get wafers for less than $100k.

D’Angelo: So if I may, I have my last comment. He said at some point that he felt the sensor was not correct. And I still feel the sensor is not correct! I hate that I have zero hardware background, but I would love to see some lateral connections in new event-based cameras. From the researcher’s point of view, it would be nice to have a sensor that has some lateral connections and does something at the sensor level.

Etienne-Cummings: Go back to Kareem Zaghloul and Kwabena Boahen—they did the first real event-based sensor that was actually modeled on the retina and so on, where they had the temporal dynamics as well as the sustained cells. Those had massive lateral connections. They took advantage of diffusers to create networks that were very similar to their horizontal cells and so on. So it’s just that, along the way, we’ve gotten much more desirous of very clean pixels that are operating individually by themselves. So then we’ve forgotten these lateral connections. But no, Kwabena and Kareem did this.

D’Angelo: Let’s stop it here for today. So again, thank you Sunny for this incredible interview, and Ralph for your comments as usual.

In the next episode, Sunny will talk to Patty Stabile, a professor in large-scale photonic neural networks for high-performance computing at Eindhoven University of Technology in the Netherlands. I hope you will join us then.

, ,