Vricon is on a mission to build the most accurate 3D model of the surface of our planet. They have access to what might be the largest stockpile satellite imagery in the world and are using a technology stack that focuses on very large scale image processing and in particular multiview 3D reconstruction. I this podcast interview the Vise President of Vricon, Isaac Zaworski walks us through what the process of creating the most accurate 3D model of the world looks like today, what challenges they face when process imagery and what this might look like in the future.
This episode is sponsored by HiveMapper
A platform that takes video and creates 3D mapping layers based on that data. The video can be from a variety of different sensors, does not need to be vertically looking down on the geography and each 3D output is georeferenced!
You are more than welcome to reach out to me on social media, I would love to hear from you!
In Conversation
From Race Cars to Afghanistan
Daniel: Hi Isaac, welcome to the podcast. Thanks for taking the time to do this — I realize your schedule has been pretty busy lately because you’ve just had your second child, so I appreciate that you took the time. You are the Vice President of a company called Vricon, and on the website I see a couple of words which are very eye-catching: “the globe in 3D” and “beyond Google Earth”. So already I’m thinking that you and your company have something to do with building 3D models of the earth. But before we dive into that, can you give us a bit of background about yourself and how you got into the geospatial world?
Isaac: Certainly, and first thank you very much for having me on, it’s a pleasure to join you. I always enjoy getting a chance to talk about this. My background is a little bit strange and I’ve taken a path that I don’t think many of my peers have. I actually started my career as a mechanical engineer building race cars. It was only out of pure happenstance that I responded to a Craigslist ad one summer while I was in grad school, looking for an intern who thought that Google Earth was the coolest thing since sliced bread, who knew how to build supercomputers, and was a photography nerd — of which I was only moderately a photography nerd. I had never really played around with Google Earth at the time, but I had been playing around with supercomputers since I was 15.
Isaac: So it just so happened that responding to that Craigslist ad got me a job at one of the only startup companies in Portland, Oregon that was doing ISR related work for the Department of Defense. So I promptly went from being a grad student in upstate New York to deploying to Afghanistan in support of an airborne ISR program, right in the middle of the surge as it was happening back in 2010, and promptly got hooked on the geospatial mission and on the broader national security and defense mission as a whole.
Daniel: Firstly I want to say congratulations on having by far the most interesting story of all the people that I’ve interviewed on the podcast so far. Going from being a mechanical engineer to working in geospatial, and especially on the path that you took — there are very few people that have a story like that. So we understand a little bit about your background now. What is it that you do?
Isaac: That’s a great question. I get asked that question on a daily basis, and I affectionately describe myself as being the Vice President of Stuff, because I play an interesting in-between role. I spend the majority of my time bridging the gap between the end users and our ultimate customers, along with the procurers — navigating the bureaucracy of a lot of the large government and commercial organizations that we do business with — but then also managing the direct interaction and shaping of our product development path and our strategy.
What Vricon Is Building
Daniel: Perhaps we’ve dived in a little deep here at the start. Perhaps you could tell us about the problems that Vricon is solving.
Isaac: That’s probably a good place to start. The one sentence description of what we’re trying to do here at Vricon is: ultimately we’re trying to build the highest resolution, most accurate 3D map of the face of the planet, and then enable new uses for 3D geospatial content across a broad set of different industry verticals.
Daniel: So if we take the first bit of that sentence — we’re trying to build the most accurate 3D model of the planet in the world. How are you doing that? What data are you using, where is it coming from, what does the process look like?
Isaac: We’re a little bit of an interesting company to begin with. We are only four and a half years old at this point in time, but we are the product of a joint venture between two very large companies, so you can think of the two of them as investors in us. The first one is a company called Saab — many of you are familiar with their cars, but in fact they are a relatively large industrial organization from Sweden that does everything from building fighter jets and shoulder-fired rockets for the United States military to building submarines, and pretty much everything in between. And then our other investor was the company formerly known as DigitalGlobe — now they fall under the larger organization Maxar — but they are the world’s leading commercial high-resolution satellite imagery company.
Isaac: When those two companies came together they each invested in us. On the one hand, Saab invested a technology stack out of their portfolio and some intellectual property focused on very large scale automated image processing, and in particular multi-view 3D reconstruction. On the other hand we had DigitalGlobe at the time, that invested all of their commercial satellite imagery. So essentially we have a live library card that gives us access to their entire commercial imagery archive — past, present and future. Today that archive sits at right around 100 petabytes and growing very quickly, worth of the most accurate, highest resolution commercial satellite imagery that’s available.
Isaac: So with those two building blocks, when Vricon was formed in 2015 the number one laser-focused goal for the company was to be able to bring together those two investments and establish a commercial production capability here in the northern Virginia area, where we could scale out those image processing algorithms that we got from Saab, and build up the workflow and the infrastructure to allow us to take pretty much all of that imagery from DigitalGlobe — and ultimately as much imagery as we can get from anyone else in the entire world — and push it all through a giant supercomputer. Automatically process it, correlate every single image against every other overlapping image, and then use that information in order to reconstruct the photorealistic 3D world.
From Stereo to Multi-View Photogrammetry
Daniel: When we’re reconstructing the world in this way, I’m assuming that we’re not just making a ball floating in space that is made up of overlapping images. Is that correct?
Isaac: That’s correct. I’ll try to stay reasonably high level from a technical description, so if I fall too far down please pull me back. The concept of stereo photogrammetry has been a core part of geospatial science for 70 or 80 years now. The basic concept is that you can use image processing techniques to derive depth from two overlapping images that were collected from different perspectives, in the same way that our eyes do. Your eyes are sensing the same scene but from slightly different perspectives, and your brain, many times per second every day, is measuring the slight variations in the location of each feature in the scene based on those two different perspectives — and from that variation your brain can measure the depth of the scene.
Isaac: So stereo photogrammetry is basically that same concept, just applied to images collected by sensors. That was the state of the art through the latter half of the 20th century. A lot of the elevation models, the large geospatial datasets that describe the terrain of the world, have been derived from stereo photographic methods one way or another. Then ultimately, as computer vision algorithms and modern computational capability really started to take hold over the last 20 or 30 years, you saw the academic community transition from that concept of stereo photogrammetry to what was then called multi-view stereo photogrammetry.
Isaac: The biggest difference there is that with traditional stereo you think of having two images, just like your eyes, and those two images are expected to be collected from a very structured geometry, so you know exactly what the angles are between them. The idea with multi-view was to remove some of that sensitivity or dependency on collection. Rather than just thinking about having two images, now I just say give me a big pile of overlapping imagery that may have been collected from a myriad of different perspectives, maybe over some period of time, maybe from a bunch of different sensors — and I will use computer vision algorithms to take each image and every other image that it overlaps with and very accurately correlate them together, so that I can then essentially backwards calculate where the original sensors were in space and calculate those measurements to understand the depth, or the three-dimensional component, in every scene. And then you get to add in this computational statistics component where you can do large scale bundle adjustments and statistical correlations through the entire stack of stereo pairs that you’re generating over any given point on the earth, and it allows you to create a very accurate end result.
Daniel: And what is the end result of all this?
Isaac: The rawest product that we at Vricon produce is what’s called a fully textured 3D surface model. In our case we’re a little bit unique — there are not that many groups out there in the community who have taken the approach that we have — but our very first product is what’s called a textured mesh. You can think of it as the type of data that you would see in a video game like Call of Duty or something like that, where you have these maps that are made up of a bunch of triangles with textures mapped to them that represent a continuous surface in the scene.
Isaac: So for us, the first thing that comes out of our automated production process is this very dense textured mesh that represents the reflective surface of the earth in a continuous way. When I say reflective surface, I mean we’re working from EO collection, for the most part from a satellite standpoint, so we can only model, from a mathematical standpoint, what the satellites can actually see at any given point in time. So we end up with this 3D model of the exterior of everything you can see from space.
Clouds, Water and the Hard Problems
Daniel: So you said that you make a couple of products and this was the first one that comes out of the system. I’m assuming even with this first product you’ll run into a few difficulties around cloud cover, for example.
Isaac: There are a lot of complexities in any form of image processing, and anybody in the field that you talk to will give you an “it depends” answer at some point in time. One of the really unique things for us is that we are trying to tackle this image processing problem at a global level. If you go out into the commercial marketplace today you’ll find a reasonably competitive landscape for desktop multi-view stereo reconstruction software packages. There are many of them — the academic community published some open source libraries a number of years ago that has spurred an interesting market. But most of those companies are focused on small scale drone-collected or airborne-collected datasets, where you can choose how you do your collection so that you don’t have to deal with things like atmospheric effects or clouds or the temporal changes that you get from bigger datasets.
Isaac: So we’re really kind of unique in the marketplace in that we actually have access to this massive amount of satellite imagery, which has both given us the opportunity and also forced us to develop our algorithms and our workflows to be able to address some of these really complex challenges that come down to the fact that environmental effects are hard, clouds are hard, and the earth is a really complex place.
Isaac: Clouds, it turns out, are not the hardest problem that we typically face. We have implemented a pre-processing step in our overall workflow where we can go through and pretty accurately mask out all of the actual cloud pixels using an automated algorithm. The place where it gets a little bit more challenging with clouds is actually the shadows that come from clouds. Because when you start wanting to use the pixels from these images to actually texture this 3D model — to add the image information to it — and you’re doing these really complex composites of a number of different images, trying to separate out the areas that have been shadowed from areas that are not shadowed, so you get a nice clean colour balanced texture, is a pretty hard problem.
Isaac: But when it comes to the 3D reconstruction side of things, the biggest problem that we see without a doubt is water. Most people who do work in the geospatial world, or who have ever dealt with geospatial data processing, will usually point to water as really one of their biggest challenges. That’s just because it’s really complex and it moves. When we’re talking about wanting to build a 3D model of an entire country, where you have say a 10 meter tidal shift from the coast on the western side to the coast on the eastern side, plus you have river networks that change elevation by hundreds if not thousands of meters across the entire terrain of the country — trying to figure out how to automatically model all of that complexity into a finished static product is a perpetually challenging problem for us.
Updating a Model of the Whole Planet
Daniel: You’re doing this at a global scale, and these data products that you are making, I am assuming they need to be maintained — they’re not a static thing at all. I’m assuming you don’t just create one model and say that’s it, we’re done. I’m assuming this process needs to be repeated on perhaps a daily, weekly or monthly cycle.
Isaac: Probably the second most common question that I get from customers and users comes down to: okay, well, what’s the update schedule, what’s the update rate? And this is one of those really tough questions when you start to think about it. Our goal number one is to go out and build a foundational dataset for the planet, and it has taken 10 years’ worth of collection for DigitalGlobe to get an archive worth of imagery that has even enabled us to push towards that vision. The reality is that even though we have access to all their imagery, and we have relationships with many of the other commercial satellite providers that are out there and use their imagery as well — the reality is that we don’t control collection. When it comes to the question of how frequently things are updated, we are completely at the mercy of when new data becomes available, when we’re talking about this large commercial production scale.
Isaac: So we’ve taken a few different approaches to thinking about how updates occur. On the one hand you can say, well, most of the collection that occurs on a daily basis is driven by priorities set by the biggest customers of these commercial satellite companies, which means that most of the data being collected is to satisfy either national security concerns or problems dealing with economically important questions. So they’re taking pictures of areas where people are, for the most part, or areas where we’re concerned from a national security standpoint. Which actually satisfies a lot of the need when it comes to updates, because it’s not a question of updating the entire planet on a daily basis. Maybe someday that will be something that we can talk about, but today that still counts as computationally expensive to the point of not really being tractable.
Isaac: But you can start thinking about trying to identify the places that are meaningful to update, and that’s one of the biggest areas of focus for us in the near future, and with a couple of our customers. Can we be smart about how we identify areas in our model where meaningful change is occurring, and then start coordinating tasking and updating — whether it be from satellites or whether it be from other sensors that may be available, so airborne or drone collected sensors — and really just have a mathematical framework that allows us to register any of that new content into this foundational model and fuse it all together?
High Frequency and Low Frequency Noise
Daniel: I’m imagining change detection would be a big part of what you’re doing in terms of figuring out what areas need to be updated. So where has change happened, and how do you do that? I’m assuming you have some algorithm running against the backlog of these images, or perhaps against the model, and then comparing it to new images that are coming in — and then I guess figuring out whether it’s a meaningful change. This must be difficult.
Isaac: Yes. Many tens of billions of dollars have been spent on trying to tackle the change detection question over the years, and you hit the nail on the head: identifying change is typically not the problem. Being able to filter out and identify meaningful change is the really hard problem. Internally in our own production process we have some tools that enable us to get a better idea of where meaningful change might be happening than many other customers, and some of that comes down to this whole concept of how the metadata from these stereo correlations can be used. If you think about it, we’re taking all of these images and we’re comparing them against the entire historical record — all of the other images that have ever been collected over a given location — and we’re doing these very accurate, pixel level comparisons of those images. So from a metadata standpoint you end up with some interesting information. Some of it comes in the form of things that we have to try to deal with or mitigate for from an initial production standpoint, but it can be quite instructive from an analytic standpoint when you’re trying to understand what the world is actually doing.
Isaac: I typically try to describe, in a simplistic way, the fact that we see two kinds of noise from an algorithmic standpoint in our production process. You have high frequency noise and low frequency noise, and these relate to how the world is changing across all of those images that we’re looking at. High frequency noise you can think of as things that move quickly in a scene with respect to the time scale of satellite collection — so things that move from day to day. If you look at 3D models that we’ve built of Washington DC, for example, where I live, you can zoom down and you can see the 395 freeway, and there will be no cars on our 3D model of the freeway. The reason is not because there were no cars on the freeway when the images were collected — it’s because there were cars on the freeway every day, in every image that was collected, but they were different cars in slightly different locations. So when you try to match a pixel that represents part of a car from one image to all of those other images that we’re comparing it against, there’s essentially zero correlation. So the algorithm says, this is high frequency noise, this thing doesn’t really belong, it’s not part of the static scene, so I’m going to choose to remove it from the finished model.
Isaac: The other type of noise that we tend to see is more of a low frequency noise, and this you can think of as things that change slowly over time. These are the types of meaningful changes that are probably more interesting in most use cases, at least when it comes to geospatial mapping. These are things that can range all the way from something as simple as agricultural fields — this was an early one that we had to figure out how to deal with, but it turns out that the elevation of corn fields changes pretty significantly during the course of growing and harvest season. So you get this situation where the algorithms are trying to model a three meter change in elevation over time, and it is plus three meters over a kind of steady increase and then a very fast step function minus three meters. That can create some confusion from the algorithm standpoint, and we’ve had to figure out ways to essentially identify that these are crop fields and apply more of an optimized algorithm for that scenario.
Isaac: But you can also think about buildings being built. If you’re trying to build cities in China, this is a particularly acute problem, where you have new skyscrapers that are getting built over the course of six months. So we may have satellite images where in the first image that we’re using there’s just a hole in the ground, and in the newest image that we’re using in our model there’s an entirely new skyscraper there, and the rest of the images represent something in between. Internal to our process that ends up as this correlation curve where you have lower correlation as you move up the building — because maybe the ground areas are represented in all the images and the top of the building is only in one image. You have this curve of the correlation of these stereo pairs that says, I’m really confident that the ground is here, and I am not at all confident that the top of the building is here, and everything else is somewhere in between. From keeping a scalable production environment, that’s a tough challenge that we have to figure out how to address, because obviously we want to be able to represent the building as best we can. But from a change detection perspective that sort of correlation metadata gives us some really interesting indicators to look for.
Isaac: With all that being said, I will also include the fact that in order to really get to the vision that we’re trying to execute on, we completely recognize that we need to be able to leverage other remotely sensed data and other phenomenologies as inputs into this larger view of the world, in order to have any chance of really getting a meaningful change detection capability that can inform both our production and also the customer’s collection strategies.
Why 3D? The Answer Is Accuracy
Daniel: I’d like to shift gears a little bit and talk about the use cases for this, because I think by now the listeners understand that you’re building this incredible dataset, this incredibly accurate and updated model of the world. What are people using it for?
Isaac: That’s the more fun part, at least for me. I joined before the company was even the company, so I’ve been with it from before its current life. It was amazing — even internally inside of the company for the first couple of years we would sit around our tables trying to strategize about how to go to market and how to get customer adoption, and one of the questions that I would always ask is, why 3D? Because in our own internal discussions it just seems like such a no-brainer. This is clearly — everybody must need this if you look at it. 3D geospatial data is one of these very visceral things: if you start looking at a place and really start exploring the world, and understanding that you’re being immersed in a part of the world that you never thought that you could see in this way — if you’re a geospatial person it’s a no-brainer. But translating that into customer requirements and specific use cases was challenging up front.
Isaac: So that question of why 3D — what is the requirement that actually drives the transition from 2D or traditional mapping techniques into the third dimension? Ultimately, if I want to boil all of the answers down to the simplest one that we have come up with, and that has really taken hold amongst all of our customers, the answer is accuracy.
Isaac: The reality is, if we all look outside of our windows right now, we have to acknowledge that we live in a three-dimensional world. And for those of us who live in urban areas, that third dimension is a critical component of our daily lives. You go to your office building and you take the elevator up or down a number of floors, and when you think about all of the other people in the building, there are a thousand other people who do the same behavior but they may not go quite as far up or down.
Isaac: As technology has progressed over the last 10 or 15 years and remote sensing capability has just exploded to the point of almost becoming ubiquitous in our ability to collect data about the world, we are increasingly — both in the commercial world and also on the government national defense side — driving the questions that we’re trying to answer down to a human scale. So we want to understand human behavior at the individual level in the world. And it turns out that when you want to understand human behavior at the individual level, you need a data structure that is capable of representing the world at human scale — which is to say that third dimension becomes a critical requirement. Because if you’re trying to figure out how to accurately correlate a number of different datasets collected from different sources and different places, where the common focus is a person, then you need to be able to represent the fact that that building has multiple floors, and that person is not on the ground floor, they’re on the sixth floor.
Isaac: You can think of a lot of the focus on the modernization of the 911 system. A big part of this is a recognition that urbanization of our communities has left an information gap for first responders and their ability to know where an emergency happens in the third dimension. Because you get people calling on VoIP phones where you don’t necessarily have a direct ability to triangulate a location, because you’re running off of an IP address. And all of a sudden you’re trying to send emergency services to someplace that is in this lat-long location, but that’s no longer precise enough for what we’re trying to accomplish.
The Gap, and Fusing What We Already Have
Daniel: I’d like to talk a little bit about the future. When I think about data collection at this scale, and I think about our increasing access to data from these different sensors, I wonder sometimes: is the future just more of the same? Can we just expect to get more data and more pixels? Obviously that’ll lead to a lot of opportunity in terms of derived data products, but is that the future itself — just more of what we have today, a little higher quality or a little higher resolution? Or are we going to shift and go in a completely new direction?
Isaac: That’s a good question, and I’m one of those types of people who is rarely satisfied with the way things are, or with incremental change — particularly when I spend a lot of my time dealing in the depths of a lot of these bureaucratic organizations where incremental change from where the users are today is still 20 years behind where the bleeding edge of technology is. It’s a reality of large organizations. The US government is a very large organization, and from an investment standpoint it takes a long time for things to change. But I’m perpetually the one who is pushing for meaningful change in order to really make a difference.
Isaac: So I look at a few things when it comes to the question of what does the future hold. There’s a technological component to this where the advancements in remote sensing probably will be incremental. Those increments have been relatively large over the last 10 or 15 years, driven by a lot of the major commercial markets that are out there, and I think there’s going to be continued growth in our ability to sense the world. That growth will be better resolution, it’ll be more spectral bands — a lot of the investment in multispectral and hyperspectral, and things like the recognition that there is information that’s outside of the visible band that we typically perceive the world in. That’s going to continue, and also looking into the RF and cyber domains as well.
Isaac: But I think the real leap-ahead opportunities are going to exist in the integration and fusion of those datasets together. Certainly in the national security world the topic of data fusion has been one that has been buzzwordy for 15 or 20 years now, maybe longer than that, and we have never really done it well — kind of like change detection, it’s never really been done well. I think we’re getting to this point in the technology space where really looking at data fusion and data integration, and at the very base level the correlation of all of these different types of data, is the major leap ahead in our ability to derive insights and information from raw datasets.
Isaac: The imagery world is an easy one to look at as an example. Even today the majority of the information that is generated from imagery that is being collected comes from people — which is amazing. There are still today, in 2019, massive organizations of people who go in and look at every picture that we’re collecting. And that’s just pictures, that’s just visible images — that’s not thinking about how those pictures correlate with, or can be analyzed in concert with, other collection that is happening over those same locations at the same time in order to generate more pertinent information. There’s a lot of work being done in fields like entity resolution, where you start thinking about how do we take these very massive and ever growing datasets, across phenomenologies, across the entire portfolio that you have available to you, and enable smart ways to use each of those measurements to inform a larger information picture that isn’t just tied to a single data type. There are a lot of really challenging problems in there — there are data problems, there are processing problems, there are opportunities for machine learning and deep learning to make big impacts, although there’s still a lot of core science underneath that as well that needs to happen. There’s still a lot of learning from the way that we humans digest and interpret data that is important to take into account there as well.
Daniel: I’d just like to take a minute to try and summarize and pick out some of the really important observations I heard you make. One of them was this recognition that there is a huge gap in the industry — it sounds like we have the people on the ground floor and the ones at the cutting edge of the geospatial industry, and there’s a massive gap between them. And I also heard you say that perhaps the next big opportunity is not so much collecting more, it’s perhaps learning how to use what we have already and integrating these different datasets. Is that correct?
Isaac: I think so. I’ll expand upon both of those points just a little bit. As far as the gap that exists, I actually think it’s a gap between not just the geospatial people on the ground floor, but between everybody — all of the users — and the people who are on the cutting edge of the geospatial world. I have spent the last several years supporting the United States Geospatial Intelligence Foundation, and one of the biggest educational and policy level topics and discussions that we have is that geospatial as a discipline is pretty niche still. Even geography is pretty niche. I will be the first one to admit that when I think about geography I think about the pull-down maps in fourth grade, and that was about all that I thought about — it was the last time that I could remember in my academic career thinking about geography as a thing.
Isaac: But the reality is that every single person in our country and in the developed world is a user of geospatial technology, and they just don’t know it, or don’t have a language to describe it. Every day we’ve become so reliant on GPS and Google Maps and Yelp and all of these other things that we use to navigate our daily lives, and all of that comes back to this geospatial data science underneath, and people are not terribly aware. So there’s a communication gap that I run into pretty frequently, from users — the people who have a need who don’t know how to describe that need in terms of geospatial capability, or what’s available, or the technology that could solve the problem. That’s a big component of that gap. And then obviously there are bureaucratic and organizational components to that gap as well, as far as friction to bringing in new capabilities, new technologies, new data, new enterprise architectures. So I think that gap is really critical, and it’s a pretty massive gap, which provides good opportunity.
Isaac: And the second point, around the opportunity perhaps not being creating more data but being better at using the data we have — that one I think hits the nail on the head. Vricon as a company is a shining case study, an example, to demonstrate that there is massive opportunity in figuring out how to use existing datasets in better ways. Our entire business plan and business strategy was built around the fact that this data has been collected and it’s sitting in data centers and in archives around the world, and if we figure out a smart way to bring it in, process it in a unique way and generate additional derived information, we can create massive value on top of what has already been collected, or what will already be collected.
Daniel: Isaac, this has been a truly enlightening and fascinating conversation for me. I really appreciate you taking the time to teach us all a little bit more about what you do and the geospatial industry in general. Before I let you go, can you let the listeners know where they can go to learn more about you, or where they can go to reach out if they have questions?
Isaac: First off, thank you very much for giving me an opportunity to come on. This is one of those things where you can probably tell I’m reasonably passionate about what I do, and I thoroughly enjoy getting an opportunity to share with the world. For those of you who are interested in more information you can visit our website, vricon.com, and if you have interest or thoughts or questions there’s an info link on our website that you can reach out through directly, and it will probably come to me directly.
Daniel: Isaac, thank you so much for taking the time, it’s much appreciated.