Settings

Theme

Linux containers in 500 lines of code (2016)

blog.lizzie.io

89 points by mkornaukhov · 22 comments

Reader

6 threads
js2

(2016). Previous submissions w/comments:

https://news.ycombinator.com/item?id=30623372 (250 points | March 10, 2022 | 27 comments)

https://news.ycombinator.com/item?id=22232705 (267 points | Feb 4, 2020 | 29 comments)

https://news.ycombinator.com/item?id=15608435 (440 points | Nov 2, 2017 | 53 comments)

abidinberkay

This was written in 2016. What would be different if you wrote it today? For example would cgroups v2 or newer seccomp features change that much?

smashed

Coincidentally I mis-prompted claude code the other day while working on a toy project and failed to specify the project should be built on top of docker and not "like docker".

It went on to waste all my tokens creating a specialized docker clone. Cool I guess.

setheron

I have written https://fzakaria.com/2020/05/31/containers-from-first-princi... a while ago in similar vein.

kragen

Last night I was looking for how to run Graphviz on untrusted input in a secure way, because recent versions of Graphviz give untrusted input to a whole insane rat's nest of code: Harfbuzz, Pango, libfribidi, libthai, libgraphite2, and its own internal format parser, each of which has a rap sheet of CVEs that makes Charlie Manson look like a petty shoplifter. And apparently Pango is even multithreaded, so we can expect nondeterminism. (Most of this doesn't show up in a simple ldd check; Graphviz sneakily waits until runtime to dlopen graphviz/libgvplugin_pango.so.6!) So, naturally, I wanted to sandbox it so that the worst thing a malicious attacker could do would be to make it draw Dickbutt or something. What I ended up with was less than 500 lines of code using Claude's suggestion of Bubblewrap http://canonical.org/~kragen/sw/dev3/wrapdot:

    #!/bin/sh
    # Confine dot in bubblewrap, taking input from stdin and writing PNG
    # output to stdout.

    # 64 megs seems to be enough, 21 megs isn’t.
    address_space=64001000

    # With zero --fsize, we can’t write the output file on stdout if it's
    # redirected to a file, but you can pipe it to `cat`.
    file_size=0

    cpu_seconds=5

    # Apparently Pango or fontconfig is multithreaded now‽
    # (process:2): GLib-ERROR **: 00:37:18.600: creating thread '[pango] FcInit': Error creating thread: Resource temporarily unavailable
    processes=4

    # We’re using --unshare-user, etc., explicitly, because --unshare-all
    # uses the wimpy --unshare-user-try and --unshare-cgroup-try options.
    # --remount-ro / prevents malicious code from filling the root
    # filesystem with empty files.

    exec bwrap \
          --ro-bind /bin /bin \
          --ro-bind /lib /lib \
          --ro-bind /lib64 /lib64 \
          --ro-bind /sbin /sbin \
          --ro-bind /usr/lib /usr/lib \
          --ro-bind /usr/share/fonts /usr/share/fonts \
          --ro-bind /var/cache/fontconfig /var/cache/fontconfig \
          --ro-bind /etc/fonts /etc/fonts \
          --remount-ro / \
          --unshare-user --unshare-ipc --unshare-pid --unshare-net --unshare-uts \
          --unshare-cgroup --die-with-parent --new-session --cap-drop ALL \
          --clearenv --setenv PATH /bin \
          prlimit --as="$address_space" --fsize="$file_size" \
                  --cpu="$cpu_seconds" --nproc="$processes" \
                      dot -Tpng -Gdpi=192

    # For testing, to verify that network access is indeed blocked:
    #                 nc.traditional -v -v 127.0.0.1 8000
Still, this is enough code that I'm not sure I haven't left something out. Still pending: run ImageMagick or netpbm inside the sandbox to convert the PNG file into a PPM or BMP — there have been CVEs in libpng in the past, and of course it's potentially vulnerable to zip bombs.

Of course this is still exposing most of the Linux kernel system call interface, although fortunately not /proc and /dev. And it's still fucking insane that drawing a node-link graph with three nodes requires more virtual memory than my first Linux machine had in total, in which it ran web browsers and recompiled the kernel. But that's a little further down the line.

ranger_danger

> I wanted specifically to find a minimal set of restrictions to run untrusted code.

I don't think we should consider containers to be a security boundary. Even full VMs can be escaped, and have been, many times.

The fact that this is possible in the first place makes me think we need a much better approach.

  • chubot

    As far as I know, Firecracker, gVisor, and Kata Containers are the solution here. They use VM primitives (x64_64 and ARM64 extensions) and have lighter codebases

    https://firecracker-microvm.github.io/

    https://gvisor.dev/

    https://katacontainers.io/

    But I don't have any direct experience with any of them. I'd be curious what people who have built on top of them think

    edit: OK it looks like Kata can use Firecracker, so as far as isolation, it's either Firecracker or gVisor. And Firecracker is the VMM I mentioned, but gVisor is quite different -- it's more like a user space kernel that emulates syscalls.

    • binsquare

      I'm going to toss in smolvm as well because firecracker needs some expertise to make the box usable and secure.

      https://github.com/smol-machines/smolvm

    • johnsmith1840

      I've deploy gvisor, done basic test of firecracker and an honest attempt at production kata.

      Firecracker and gvisor are nice systems not horrible to use, gvisor isn't quite the same security level though.

      Kata is HARD to make. The technical know how to make that in production is awe inspiring. I wanted to use it but it was so complicated to integrate into a cluster I literally just gave up and mirrored raw VMs into the cluster which was alot easier actually.

      Kata also breaks any potential of confidential VM unless you're a virtualization wizard.

      You should go check out redhat's confidential container method for a production design overview. Their ARO self hosted system.

      • palata

        > gvisor isn't quite the same security level though

        Which one is more secure? I thought gvisor but your sentence sounds like it implies the opposite.

    • laurencerowe

      As I understand it Kata supports multiple VMM backends, Firecracker, QEmu, Cloud Hypervisor, and their own Dragonball. Except QEmu, I believe those are all built on crates in the rust-vmm ecosystem, each making slightly different tradeoffs.

  • zamadatix

    We don't have any "security boundaries" by this definition, just "security make-it-harder"s. I.e. "security boundaries" always have a relative strength associated with them, not a guarantee they keep the thing secure without any doubts.

    • SOLAR_FIELDS

      The principle of defense in depth is built around the idea that with enough time, any system can likely be compromised, but the chance of compromising the system before being detected in your attempts to do so is lower the more safeguards you put into place

  • bityard

    Depends on the threat model. Security is not black-and-white.

    Containers protect against "I don't trust this curlpipe to not crap all over my dotfiles," rather than, "there might be a sandbox escape attack in this random file I downloaded."

    If a VM is not sufficient for your threat model, I'm curious what is?

  • stryan

    Podman supports using KVM backed virtualization for containers via libkrun: `podman run --runtime=krun` . Still not the end-all-be-all security boundary, but better I think.

  • raesene9

    I definitely wouldn't trust standard Linux style containers that expose a shared Linux kernel at the moment, there's been far too many LPE and container breakout vulnerabilities this year. It's possible that in future if the kernel gets a lot more hardened, that could change but things like Firecracker are a better bet from a security standpoint.

  • LtWorf

    They are a security boundary, but like everything else, not perfect.

  • mdspan

    I think until something hardware-based like CHERI becomes widely deployed (which seems extremely unlikely in the near to mid term given), we're going to keep seeing VM escape CVEs pop up indefinitely.

    • LtWorf

      Because we've never encountered hardware bugs…

      • mdspan

        Critical hardware bugs occur an order of magnitude less frequently than critical hypervisor/kernel bugs, which is why they always make the news. In general, they're also more difficult to exploit. We haven't seen any serious spectre or meltdown malware in the wild almost a decade later.

Keyboard Shortcuts

j
Next item
k
Previous item
o / Enter
Open selected item
?
Show this help
Esc
Close modal / clear selection