Unikernels were hard. key word: were
ghuntley.comAs mentioned before on this topic. What about debug ability? An application overflow now corrupts part of the network stack.
In an oxide episode there were some mentions of reading off data lines but I just don’t think that’s practical.
The reduced attack service is cool but not at the expense of my visibility and liveness of the system
Yeah that's generally the argument for using a managed runtime (like Ocaml in MirageOS's case, or I could see Go or even the JVM fitting here) for these kinds of things. Running a VM which has no pointers, memory access etc primitives, and is garbage collected etc direct on "metal" gives more peace of mind about that sort of thing.
You could make the argument that Rust w/ its memory safety is a candidate, but w/ Rust it's still entirely possible and fairly easy to break out of that. And Rust/Cargo applications have a habit of using a bazillion third party deps that you then need to keep a close eye on.
Unikernels are on the smaller side of the spectrum; Linux on the larger side, with a lot of room in the middle...
For that matter, the actual Linux kernel need not be large at all. It's really the user space stuff that takes up the bulk of what people think of as the largess of Linux.
Still, there's a whole pile of stuff there that many applications don't need. If you truly can run the entire of the OCaml runtime and standard libraries etc without it, and you're happy writing your apps in OCaml... MirageOS has always felt winning-ish to me as a general idea.
And to TFA's point, now with LLMs the relative exotic nature of OCaml as a platform need not be a huge hindrance. I like the language, but have only dipped my toes in. Definitely tempting to try writing apps on it w/ MirageOS now. Though I doubt I'd convince anybody to pay me to do that...
Reducing attack surface is definitely a plus but it is nowhere close to the number one security benefit of running unikernels.
That's why I never really liked talking about "reducing attack surface" that much because folk inevitably turn to lines and code, which while reducing is good, just simply doesn't communicate what the biggest problem truly is.
Vuln exploitation is the number one entry point for data breaches and os command injection is the number one CWE in CISA Kev from last year.
System intrusion was repeated something like 64 times in last year's DBIR.
The operating system itself is literally the problem as it's inherently meant to run many different programs whereas unikernels only run one.
> The operating system itself is literally the problem as it's inherently meant to run many different programs whereas unikernels only run one.
Unless you rely on security through compartmentalization. See: https://qubes-os.org
I’m completely ignorant about this and likely just missing the point but: isn’t the point of an OS to not have to (vibe)code filesystems and networking by hand every time you need them? Also, how does a unikernel cooperate with other applications? Would they all live in separate networked unikernels managed by a hypervisor? And, if they all (vibe)coded their own fs and network wouldn’t that introduce subtle bugs and inconsistencies that would eventually bring the whole thing down and lead you back to the need for shared primitives in the first place?
Maybe a good halfway point is to still have a unikernel but the libraries for common stuff like networking or fs are already written according to a standard and plug and play and reusable across applications?
I guess the others have made these points, but since unikernels run in VMs with virtual hardware, I guess a virtual ethernet card can have an interface that's easy to program for, making the driver trivial. You're not talking to real hardware in any case.
Similarly, a filesystem's just a data structure that happens to be on disc. You can mmap your harddrive and let the host hypervisor handle swap, preemption etc.
It's similar to a process contract, but with slightly different boundaries
Unikernels simply redraw abstraction boundaries. You can still use libraries and third-party code to implement common primitives. In fact, that's exactly what many unikernel frameworks give you (unikraft, hermit, mirage)
You don't build the networking or filesystem by hand every time you need them. You pull them from reviewed and maintained library code from a repository. That's how MirageOS etc work. Someone else has written e.g. TCP/IP -- hopefully well -- and you link against it.
The point that makes this different from an OS is that there's no shared service, no syscalls to do that stuff -- just subroutine calls -- and no timesharing (except at the hypervisor level). It's just a single runtime running on hypervisor direct against the virtualized hardware.
Because, yeah, all this work has happened in hardware in the last 30 years to make that possible, but we still treat the operating system as the best unit of resource sharing. When in fact it's kind of a jack of all trades master of none. If you're booting a whole linux kernel -- with its giant framework built to co-tenant a pile of applications and users -- just to run one process, there's definitely something to look at there. Even if the unikernel space itself is relatively immature.
You don't need to emphasize the "vibe coded every time you need them" thing. That's not what anybody was being talked about in the article. The "vibe coding" piece is pointing out that the missing pieces in the ecosystem can be more rapidly filled in now because of agentic tooling.
> running on hypervisor direct against the virtualized hardware.
Well, here is your actual OS then. You named it "hypervisor", and changed the system interface, but it's still there, revirtualizing the access to hardware and ensuring separation of different users/applications/unikernels.
This isn't too dissimilar to DOS programs. DOS also gave you a filesystem, memory management, a shell a process loder, driver interface, but it was all in real mode, and you were free to replace these components with your own as you pleased.
Or if you're more familiar with embedded, like a BSP or RTOS
I mean, sure we can put labels wherever we want. I don't really care. MirageOS even has "OS" in its name.
The point is all the other things: total number of lines of code, permissions and security, memory management, drivers, process mgmt, it all changes.
Does it though? And I'd really like to see how you'll revirtualize e.g. an NVIDIA GPU (notoriously context-full piece of hardware), while allowing transparent shared access to it from different applications. Heck, you'll run into trouble with most of PCI-connected devices except for the most primitive once.
yea, "virtualizing" an NVIDIA GPU is not truly feasible right now from what I see.
I do know that AMD has done more in this space. When I worked at Google there were (other) people on the Stadia team doing this. And some open source bits out there that I've seen, including from AMD.
Unfortunately, my real world work... and most real inference etc work others are doing... is on CUDA/NVIDIA.
I do wonder what kind of wins we're going to see from unikernels.
I used to regard V8 Isolates as a best possible sort of technology, with userlands juggling lots of processes.
Seeing netlify & unikraft switch to microvm's and have such a huge speed up was a bit of an awakening for me. Those are really fast start times! https://www.netlify.com/blog/edge-functions-firecracker-micr... https://unikraft.com/customer-stories/edge-functions-netlify...
Intuitively, I think I have some appreciation for how much silicon has been poured into virtualization. Its always seemed like a "yeah but you could avoid those costs by not doing that" but I'm more receptive to the idea that these might in some cases be really good ways to get some of the isolation workloads demand with the hardware helping us out, these days. There's so much securing for vm's, and maybe it's just easier than trying to secure in userlands: let the hardware help.
Long time interest in microvm's but they felt heavier weight than I wanted. Now it feels like maybe they might actually in some regards in some ways be lighter weight than managing workloads yourself in userland. Maybe. I dunno. Interesting times, i'm open to it.
Edit: I just chatted with Astra Pro some about these ideas, if anyone wants some sense material to chew on. https://chatgpt.com/share/6aca924c-8e64-83ea-a5c3-aae08b2cdf...
One of the problems is there's only a small certain percentage of scenarios/applications that benefit from very short start times.
I'd wager most of the services running on the interwebs are web servers, database servers, inference servers etc that don't change over their lifetime really. They're just doing the same thing all day long every day.
If you're doing stuff like what I'm doing for work right now, which is, yeah, multiplexing potentially oodles of user-submitted jobs, and those jobs are best expressed as distinct images or containers, then yes, managing start times is absolutely imperative in improving utilization/occupancy and therefore reducing costs.
But I'm not convinced that's a typical scenario, not typical enough to drive enough time and money investment in this space maybe?
Also there ain't currently no real "hypervisor" for the (NVIDIA) GPU. Not practically anyways. And that's arguably where we need it the most. Or at least I do, for Day Job(tm).
So unikernels and microvms may have to lean on other arguments for adoption: security and simplicity-to-reason-about might be those...