Settings

Theme

I wrote an bash enumerator because I was sick of xargs

numerlab.org

178 points by wallach-game 19 hours ago · 172 comments

Reader

account42 9 hours ago

> Consistent syntax — same {} placeholder for files, lines, ranges, or lists

Inconsistent syntax native bash methods so its an additional syntax to learn.

> Template mode — single-quoted commands work as shell templates: enumerate -f '*' -- 'cat {} | head -5'

Passing commands a a single string is BAD. Now you have to think about escaping and quoting. What does {} get replaced with if the enumerant contains unsafe characters? Can it be used as part of a larger argument or only on its own? Who knows, it's not bash. Compared to a bash loop its also always a subshell with all the implications that has - even find can be piped into a normal bash loop.

> Filters — --include and --exclude with glob patterns

A fraction of what find or native loops provide. And since this can't replace them in general enumerate is an additional thing to learn on top.

> Extensible — drop a file in lib/enumerators/ to add custom sources

To extend bash loops you don't even need root access, you just add the code to the loop.

  • iririririr 8 hours ago

    being consistent in a sea of inconsistency is a good thing, no matter how hard you try to paint otherwise. the solution have to start at some point. But the rest, yeah, scary.

zxexz 9 hours ago

I don’t want to be a downer here, but the structure of this repo and the verbosity+language of the docs feel 80% vibecoded. Not that that’s wrong; I just feel kinda gullible for even clicking on this.

One uses xargs or parallel only a few times before they remember some of the quirks that but them. And then they become cautious. And then it’s muscle memory. And if it’s not an often occurrence, they learn to check the man page.

Anyways, in a world of “vibecoding” why add another “tool” to the mix when the LLMs have been trained on all our stackoverflow-posted grievances to begin with?

giov4 6 hours ago

I lol'd, but do you think this is valid?

https://github.com/wallach-game/bashumerate/pull/1/files

to me looks like a right solution, even if I was also tired of writing for loops by hand usually

  • Dibby053 5 hours ago

    Some of these are wrong, as standard shell globs don't traverse directories. For instance the first example should be:

        ls **/*.sh
    
    I think that would work in Bash just as well as using find.

    But yeah, even in POSIX it feels redundant when we already have find, seq, grep and whatnot.

    • js2 2 hours ago

      I thought I knew every nook and cranny of bash, but somehow never noticed the globstar option added in bash 4.0.

      https://www.linuxjournal.com/content/globstar-new-bash-globb...

      https://www.gnu.org/software/bash/manual/bash.html#Pattern-M...

    • wongarsu 4 hours ago

      I frequently have cases where `ls **/*.png` fails because the argument list is too long. Just happened yesterday with an rm I wanted to run.

      Find works, but the syntax is arcane. Not bad, but unlike any other common cli tool, which makes it more difficult to remember if you haven't needed it in a while

      • jimmaswell an hour ago

        > Find works, but the syntax is arcane.

        fd is a lifesaver: https://github.com/sharkdp/fd

      • Grimeton 33 minutes ago

        Newer finds have a -delete option. Dunno if that's standard or some GNU addition, but it's there.

      • giov4 4 hours ago

        I agree and I usually ended-up combining find reliability with other commands to obtain what I needed without big issues or too much looping syntax effort.

        Regarding remembering the find syntax I think its being arcane is what it made for me more easy to remember :) I now have a unique brain area dedicated to remember only that.

    • IsTom 4 hours ago

      To be pendantic

          ls -- **/*.sh
      
      Otherwise it will fail if you've got a file named e.g. --help
      • ahoef 3 hours ago

        So, it's essentially argument injection?

        I dislike these tools so much, because you need to know all these corner-cases and I don't.

        • IsTom an hour ago

          Yeah, this happens broadly in shell with globbing. It's a bit ridiculous.

    • westurner 3 hours ago

      But `ls -print0` is not an option;

        set -x
        mkdir test; cd "$_"
        touch "test "$'\n'"12.txt"
        ls
        for name in `ls`; do echo "$name"; done
        find
        find . -exec echo {} \;
        find . -print0
        find . -print0 | xargs -n 1 -0 echo
        #
        find . -print0 | el -0 -v -x echo
        find . -print0 | el -0 -v -x echo "{0} #"
        find . -print0 | el -0 -v -x echo '"{0}" #'
        find . -print0 | el -0 -x sh -x -c 'echo "{0} #"'
        
      
      Though, I just realized that this doesn't work yet either:

        find . -print0 | el -0 -x -- sh -c 'set -x: echo "{0} #"'
        
      
      westurner/dotfiles/scripts/el: "edit lines" https://github.com/westurner/dotfiles/blob/master/scripts/el...

      `el -v/--verbose` is useful because it prints each input; though, I might not have written `el` if I had been aware of `xargs -n/--max-args=1`

  • inigyou 3 hours ago

    The main reason to use xargs is to split up argument lists before they exceed the limit.

  • giov4 5 hours ago

    On a second thought I think that the effort in trying to find a shell uncomfortability problem and trying to solve it with a "wrapped workaround" still represents positive effort and should be supported and fostered without jumping directly in 90's like rtfm mode.

    especially in the case of a younger profile, learning and putting dev effort, even if vibed or redundant, its still effort and a learning activity!

    https://github.com/wallach-game

    so I support this even if the result and problem solving can be less significant

    • Dibby053 5 hours ago

      It's possible the tool is useless and a consequence of not reading the manual, but it could just be the tool is nice but the examples are too simplistic.

      In any case, toxicity aside, publishing a command line project and having someone dismiss the examples with shell oneliners is also a valuable lesson, and I don't think it will make him a worse developer.

      • giov4 4 hours ago

        that's what I also thought when I saw the pull request, that it could be a useful lecture.

        But again the usefulness of the final tool should not compromise appreciation of effort and learning.

        Unless the tool targets, claims or pretends to be the best, final and universal solution to such problem.

        I believe it isn't, I will still fall back to my bash memories, I will not need, install and use this tool, I will not need, install or use custom shells. The effort wasted in re-aligning my brain to a tool change would superseed instantly the few second effort needed to reach the needed bash one-liner.

        Also, didn't you also feel good and amazed when completing a very long and nested working one-liner? wasn't that orgasmic? :D

dundarious 9 hours ago

If you want a shell to interact with the results, you can of course just use a (sub)shell.

    ls -1 ./*.sh | xargs -rd\\n sh -c 'for i in "$@" ; do ... ; done' sh
1. not strictly necessary to use -1 as I believe all common ls detect !isatty(stdout) and produce line-by-line output anyway.

2. xargs -r just doesn't run the command if there's no input, also not strictly necessary but I'm addicted to using it because it's the sensible default to me.

3. xargs -d\\n makes it collect fields as full lines, which is what you typically want, unless you're able to generate NULs.

4. use whatever shell you want of course, but I don't use bashisms, etc., by default, /bin/sh is fine for me, even if it's dash.

5. the trailing "sh" at the end is due to a quirk of `sh -c` usage, where $0 is the first non option argument, so `printf %s\\n 1 2 3 | xargs -rd\\n sh -c 'for i in "$@" ; do printf "%s " "$i" ; done ; echo'` (note the lack of trailing "sh") would only print "2 3 " as $0 is not included in "$@" ($0 is 1, $1 is 2, $2 is 3). It's very easy to just always give the shell name itself manually as $0 instead of trying to ingest "$0" into your logic.

Of course, you could `find . -maxdepth 1 -type f -name '*.sh' -print0 | xargs -r0 ...` instead, depending on what you're up to, that may be the easiest. It's definitely the simplest -- as long as your xargs has -0 support.

  • dspillett 6 hours ago

    > not strictly necessary to use -1 as I believe all common ls detect !isatty(stdout)

    I remember, though from the dim an distant past so it could be a long-fixed bug, this not working in at least one circumstance. I've explicitly included -1 in scripted calls to ls since. Of course we are breaking the best practice rule of not trying to parse the output of ls, so problems are not unexpected…

    As a generally thing I like to include directives like this in scripted calls even if they happen by default anyway, because is makes my intent clear: I expect the output to being in simple single-column format and the rest will break if it isn't. I even sometimes go as far as specifying --sort=name if nothing else in the pipeline is going to enforce that.

    > as long as your xargs has -0 support

    I don't think I've encountered an xargs that doesn't have this support for a long time, though I don't work with embedded stuff so maybe there are cut-down versions out there still in active use for space reasons.

    The problem I've hit numerous times is wanting to do something between “find -print0” and “xargs -0” and that something not supporting NUL as the item delimiter.

    • dundarious 2 hours ago

      I don't remember which was the odd one out, but in the 00s I was using Mac OS X and FreeBSD (and maybe some other more niche alternatives) a lot, and something didn't support xargs -0.

IdiotSavage 9 hours ago

In this thread: bash experts with arcane knowledge, unintentionally demonstrating how awful bash is.

The obvious solution would be to use something more sane, like PowerShell or nushell, but instead old experts will always defend the skills they have honed for years, while criticizing anything that's different.

  • Grimeton 26 minutes ago

    No, there are two problems here:

    1. Things have developed over time. Bash and other shells weren't always this way. If you have ever touched different shells, awk, sed and so on and then touch perl you see how it is basically just the glue between the other tools put into one language.

    2. The majority of scripts is written in sh or bash. So these shells won't go away anytime soon and any newcomer will be confronted with them. So it's best to know all the edge cases before you do stuff you don't want to do. And yes, you could always install another shell. But that brings another dependency and opens up another can of worms. In professional environments it's not really an option to have the next exotic language.

    It's like having a tmux.conf on your local machine that configures tmux exactly the way you want it. Nice to have but the moment you touch any of the other billions of systems out there that run the default configuration, you might be lost because you never learned the default and only use your modified version.

  • dundarious an hour ago

    Given the alternative proposed by OP was more of the same but with layers on top, saying "how about just using xargs and sh anyway" is a simplification -- learn 2 tiny things vs start depending on some new tool whose only purpose is to let you avoid learning those 2 tiny things by adding template language, etc., on top? I'll learn the 2.

    As for widening the scope to other tools, I've written about that on here before, how I have lots of powershell experience, some nushell, etc. I haven't become a convert (definitely not on Linux, of course occasionally on Windows).

    Happy to respond to any input or anything, but you should reply directly to people if you're going to criticize them -- very poor form to blanket criticize while also taking an uncharitable read.

  • sgarland 3 hours ago

    As others have pointed out throughout the thread, the reason we defend the skills is that there’s a high likelihood that these tools will always be available on any machine; PowerShell and nushell, not so much.

  • GuB-42 4 hours ago

    Maybe it is out of habit, but I never managed to get into PowerShell, in fact, I am not at easy with the Microsoft way of doing things, with few exceptions. Too much UNIX I guess. Nushell seems to be based on the PowerShell philosophy of using structured data and not text, not my thing.

    I really like the UNIX way of using text I/O, it has it flaws but it works for me. But that being said, I still hate bash and all its family. It has so many footguns it is an entire armory at this point, mostly related to spaces and escaping.

    Something Perl-like could be a saner replacement. It is already a bit shell-like, it doesn't struggle with escaping the way bash does, and it has very powerful text processing abilities that go well with traditional UNIX tools.

    • avisser 3 hours ago

      I'm stodgy, but I can't get past how they chose VERB-OBJECT instead of OBJECT-VERB. get-<tab> is useless. MSFT knew tab-completion would be a thing, yet they made it useless.

      • Natfan 17 minutes ago

        Get-AD* is less useless, though? modules should have prefixes! (except for Exchange/ExchangeOnlineManagement who have Get-User and Get-Group, because fuck you apparently)

  • jltsiren 6 hours ago

    You could also say that English is an awful language, because the spelling and the pronunciation have diverged too far. But it's the international lingua franca in most contexts, because it's the international lingua franca in most contexts.

    People use Bash as the default shell for scripting, because people use Bash as the default shell for scripting. If you want to replace it, you should pick a winner and discourage the use of alternatives, especially when they are better than Bash. So don't say "use something more sane, like PowerShell or nushell". Say something like "use PowerShell, or Bash if you really have to for legacy purposes, but never use nushell for any reason" instead.

  • nxobject 3 hours ago

    This is a sleeper option, but 'xonsh' as a Python superset is a very pleasant experience, especially when you just don't feel like learning another bespoke language for the complex stuff.

    I once had to scrape the output out of something hacked together by a grad student 10 years before me, scrape yet another hacked together application, do some fancy numerics, and then plonk it into a database. I did it without tears!

  • jjgreen 3 hours ago

    Bash is for whippersnappers, us old gits use POSIX sh

  • ubercore 4 hours ago

    I'm sure it's essentially a familiarity thing, but PowerShell and nushell is way too much typing for me.

  • lobofta 9 hours ago

    Switched to Nushell and I am not looking back. I don't see any major reason why we should keep dragging Bash into the twenty first century. Nushell is the first time I feel like I can write complex systems operations in a shell without having to spend either a ton of time in the docs or being at the mercy of an LLM. It is godsend.

    • piaste 6 hours ago

      I tried to switch to nushell full-time but couldn't stick with it.

      Ironically, the main friction point were not old-fashioned tools (which `from ssv` usually handled nicely) but the 'new generation' of core CLI tools like eza or fzf. They have really nice visualizations, but they do not output structured data as a middle step, so all the colours and lines only play havoc with nu's parsing.

      Since I need to "ls" a lot more often than I need to do data manipulation, the tools won and I went back to zsh. Still keep nu around for the occasional config/data file wrangling though.

  • ndsipa_pomu 4 hours ago

    Yes, Bash is filled with footguns and is awful due to that. However, it's still incredibly useful as it's everywhere and has a far greater lifespan than almost anything else. You can write scripts in PowerShell or nushell, but then find that twenty years later they're no longer usable or you find a twenty year old machine that won't have PowerShell/nushell installed.

    It's not so much about defending arcane scripting skills, but that Bash functions as a lowest common denominator and is useful because of that. If you want something that works reliably over decades, then it's best not to go for an "improved" shell as it may not still be around.

    I like to think of Bash script writing as the opposite of riding a bike - you have to relearn it almost every time you write a script.

  • kfsone 9 hours ago

    I was the kind of shell guru whose teeth itched when they saw someone doing `grep | awk`. Then I had to try and bring Windows into a Mac+Linux CI system under Jenkins groovy files that were full of `"""sh` fragments, shell scriptlets wrapping python, etc, etc, and I don't bat (I don't groovy either).

    Option 1: Learn to bat and try to translate. Hmm. No. Just no Option 2: ?

    Pwsh core had just come out. My immediate thought was that it would be great material for an anti-MS-ragging blog post, but then a line leapt out at me from one article I was glossing over: "... POSIX Terminal Shell Spec ...".

    I still wanted that anti-MS-ragging blog material, so I decided to try and use Pwsh as a Rosetta stone until I got to a point I could convert to a real language.

    But things just began to click for me. It was like going from Perl to Python - suddenly everything is an object and you can interact with everything* that way.

    There's no need for grep or awk or sed in pwsh, because the output of a shell command is an object -- a string (or []byte). It has methods.

    (netstat -an).replace("192.168.86.", "10.0.100.")

    6 years later, pwsh is the default shell on my Mac, Ubuntu boxes, lab vms, ... everything but my docker containers unless I'm feeling feisty.

    • no-name-here 8 hours ago

      I’m curious why it’s your default on most but not all platforms?

      • iririririr 8 hours ago

        ideally containers run a single process. if you're spawning shells you're doing something wrong. and idealier, you should use just bare processes with namespaces and abandon docker.

    • anthk 3 hours ago

      awk can do grep by itself.

  • amelius 8 hours ago

    Yes, Bash has no place in modern software engineering.

    The only reasons people still use it are historical and laziness.

    If it was invented today, professionals would cringe at it.

eqvinox 17 hours ago

Not sure I see the point of this. Remembering a bunch of options on one tool is no better than a bunch of tools/constructs? Especially if the latter are useful elsewhere. And it can't be used for scripts unless you start shipping it alongside, which… nah.

Just go with

  foo | while read X; do bar "$X"; done
  • xorcist 15 hours ago

    Bash while loops are pretty readable and the above would be a nice to iterate over lines if it wasn't for that gnarly pipe, which is a common source of errors in this construct. Remember that a pipe starts a new shell. So:

      grep stuff file.txt | while read key value ; do [ "$key" = "target" ] && found="$value" ; done
    
    where you might expect $found to end up with the value for the line that has "target" in the first column. Then you notice that darn pipe symbol. The variable found is set in a subshell that terminates and the value is lost. This is a problem every time you need to keep some sort of state when looping. If you can tolate a bash-ism then you could do:

      while read key value ; do [ "$key" = "target" ] && found="$value" ; done < <(grep stuff file.txt)
    
    but that doesn't read as nice and isn't compatible. It does avoid a common source of problems though, and might be worth getting into muscle memory for the times it is needed.
    • eqvinox 10 hours ago

      You can put the "< <( foo )" before the "while" btw, for pipe-like ordering of things. All redirections can be anywhere on the command line.

    • dylan604 14 hours ago

      > and isn't compatible

      is that a GNU v POSIX type of compatibility issue?

      • eqvinox 10 hours ago

        bash vs. POSIX sh (or some other shells)

        zsh should be bash compatible on this AFAIR

  • t43562 9 hours ago

    While FTW! I just wondered whether it could be simplified because I get a bit tired of typing out my own belt and braces version of it:

      { # code that generates one item/filename per line of output }  | { while read ITEM; do # do something with "$ITEM"; done; }
    
    So I just tried this and it seems to work for a trivial case although you need 1 escape:

      enum() {
        local source=$1; shift; local action=$1; shift  
        { eval "$source" ; } | { while read ITEM; do eval "$action"; done; }
      }
    
      $ enum "find /tmp" "echo found: \$ITEM"
    
      found: /tmp/.XIM-unix
      found: /tmp/.ICE-unix
  • laughing_man 15 hours ago

    Depending on the actual command, this can be far slower and less efficient than xargs. You're creating a separate process for each invocation of bar when a lot of commands will take many targets for a single invocation.

    Try this with find and grep vs xargs. There's a big difference.

    • eqvinox 10 hours ago

      3 reasons against that:

      * in a lot of cases the performance just doesn't matter

      * xargs gets you the spaces in filenames landmine

      * some commands don't even support multiple target filename arguments

      For scripts in long term use, yeah, sure, figure out xargs maybe. Any other situation with a "| while read" solution, especially on an interactive session, is an oddball and simplicity wins.

      • aa-jv 4 hours ago

        >* xargs gets you the spaces in filenames landmine

        Ignorance of xargs gets you the landmine.

        Understanding of xargs, helps you complete the mission without blowing off limbs.

            $ find . -name "Some Files With Spaces*.txt" -print0 | xargs -0 -I {} echo "Processing: {}"  # landmine avoided
      • laughing_man 10 hours ago

        xargs supports null termination.

        While it's true not every command supports multiple targets, presumably you know if you're using one of those commands.

  • JoshTriplett 15 hours ago

    I use "| while read" as well, because it works well in a pipeline and handles embedded spaces. It doesn't handle embedded newlines, but in practice, real files have embedded spaces, while embedded newlines only happen in test cases and exploits. (You can do actual NUL-delimited reads with `-d ''`, but for a quick command-line operation that's generally not necessary, and if you're going to be that careful you probably also need `-r`.)

  • aa-jv 4 hours ago

    Yes, I think the author is just showing their ignorance, not actually doing anything about that ignorance, and re-inventing the wheel:

        $ find /path -name "*.pdf" -print0 | xargs -0 -I {} echo "Processing: {}" # handles paths properly, obviates the need to write anything new whatsoever
    
    Seriously kids, learn your tools and check yourself before you wreck yourselves writing tools that really, really don't need to be written.
  • fiddlerwoaroof 15 hours ago

    I just use while’s default variable name, $REPLY most of the time. But `while` is my preferred tool for the xargs problem in most cases.

db48x 12 hours ago

> Or maybe you pipe into xargs and pray your filenames don’t have spaces…

Always use -0. Most gnu utilities support it. It makes them put a null byte after every filename instead of a newline. Completely eliminates the problem of dealing with whitespace in the filenames.

  • AdieuToLogic 12 hours ago

    To support your recommendation and redress the strawman the article postulates, the post's author could have replaced:

      find . -name '*.log' | xargs rm
    
    With:

      find . -name '*.log' -print0 | xargs -0 rm
    • tjalfi 3 hours ago

      This one doesn't need xargs.

        find . -name '*.log' -delete
      • sgarland 3 hours ago

        THANK YOU. I kept nodding along to all of the xargs comments, thinking “sure, but what about the even easier solution?”

  • dieulot 2 hours ago

    No dependency on GNU required anymore, POSIX 2024 supports it: https://blog.toast.cafe/posix2024-xcu#the-null-option

  • jsrcout 8 hours ago

    For tools that don't support -0, you can add the NULs yourself without much fuss. It's a one-liner in awk/perl/sed. Very handy.

    • ykonstant 6 hours ago

      If you do not have -0 options in your xargs/find, I wonder if you will have those extensions in your awk/sed. Perl is a different story.

      • MrDOS 4 hours ago

        I think the parent was referring to using awk/sed to do the equivalent of:

            find ... | tr '\n' '\0' | xargs -0 ...
        
        I.e., a blind replacement without the tool having any particular semantic understanding of what it's translating.

        That still requires your xargs to have -0 support, though, and I'd be surprised to have that without the corresponding option on find. But I've done this before when feeding xargs from something other than find.

        • mr_mitm 24 minutes ago

          It is also built on the assumption that filenames never contain newlines, which is wrong in general.

          It's baffling to me that POSIX allows non-printable characters in file names, but here we are. Surely they had a reason for this decision.

tester457 17 hours ago

On the topic of xargs replacements, I love gnu parallel.

The --dry-run flag of parallel made me confident to do more batch processing than I ever did with xargs.

Parallel has an option for almost everything, it's almost too much.

But I have shopped around for alternatives. The creator Ole Tange maintains a painstakingly long article of the alternatives and their differences. [0]

The gnu parallel book and reading materials [1] are excellent too.

[0] https://www.gnu.org/software/parallel/parallel_alternatives....

[1] https://www.gnu.org/software/parallel/#Tutorial

  • somat 17 hours ago

    I have used echo as a sort of poor mans equivalent for a safe check of a pipeline, removing the echo when I felt the rest of the pipeline was working correctly.

      shell stuff | xargs -n 1 -I % echo real command and % args
    
    I have also been known to write scripts where instead of executing the critical parts it prints them. Then a dry run is

      script
    
    and the real run is

      script | sh
  • dataflow 16 hours ago

    Oh man, parallel is awful. I have to rant about this because I literally tried it again today morning.

    Every single damn time I try parallel and decide to give it another chance, something ends up not working or causing a problem. I can never get it to just do what I want and get out of the way.

    Today I foolishly thought maybe I was the one who was holding it wrong every single time in the past, so I copy pasted another command that was supposed to work, and thought surely this would be straightforward. Boy was I wrong. I got some manifesto about academic citations and plagiarism, which confused the hell out of me. After I wasted time trying to figure out how to turn off that nonsense, the app just hung there trying to figure out how long its command line can be? Literally doing nothing? What the hell? I killed it but then my terminal didn't close because every time I did this apparently some perl command was spawned in the background blocked on nothing. Why the hell was perl even relevant? Nothing I wrote used Perl. Just run the darn commands I asked in parallel, is that so hard?

    • brokenmachine 9 hours ago

      I'm not new to command line tools, but I somehow managed to delete ~10k pictures while trying to use parallel to resize them.

      I forget the details but it was some kind of surprisingly weird foot-gun behavior.

      Luckily I had a backup, but it really has made me scared to try using parallel again.

    • nvme0n1p1 16 hours ago

      > the app just hung there trying to figure out how long its command line can be?

      That's common on linux. Many tools read from stdin if a file path isn't given: cat, xargs, base64, cksum, etc.

      The citation thing is a little silly, I'll give you that one.

      • dataflow 10 hours ago

        > That's common on linux. Many tools read from stdin if a file path isn't given: cat, xargs, base64, cksum, etc.

        No. I did pipe to stdin. It's not my first time using Linux...

        Here's a command line I ran right now, and the output I see:

          $ echo "http://www.example.com" | parallel -k -j 8 curl -s "{}"
          Academic tradition requires you to cite works you base your article on.
          [...more nonsense...]
          To silence this citation notice: run 'parallel --citation' once.
          
          parallel: Warning: Finding the maximal command line length. This may take up to 1 minute.
        
        So I wait a few seconds... until I get fed up and look at my process list, and I see perl is just... seemingly sitting there, doing seemingly absolutely nothing. I'm not going to waste a whole minute of my life waiting for this; I see no reason competently written software should take that long just to accomplish such a simple task where nothing is remotely close to reaching any limits.

        So I Ctrl+C. And then the parent perl process gets killed, but the child apparently keeps running.

        I press Ctrl+D to exit the terminal, and then:

          $ # (Ctrl+D pressed)
          logout
        
        ...it just sits there waiting. Ctrl+C and Ctrl+\ do nothing. I have to kill the lingering perl process manually.

        xargs Just Works without any of this nonsense, yet somehow I'm the one holding GNU parallel wrong?

        • darrenf 10 hours ago

          > parallel: Warning: Finding the maximal command line length. This may take up to 1 minute.

          > So I wait a few seconds [snip]

          This warning is only ever printed if running in Cygwin, not Linux or macOS or elsewhere. Cygwin is notoriously slow.

              # This is slow on Cygwin, so give Cygwin users a warning
              if($^O eq "cygwin" or $^O eq "msys") {
              ::warning("Finding the maximal command line length. ".
                    "This may take up to 1 minute.")
              }
          
          Also note that it only figures this out first time, after which it’s cached on disk.
          • dataflow 8 hours ago

            This is on MSYS2, yes, and that excuses absolutely nothing, because this shouldn't be happening in the first place for speed to be even relevant. At the risk of repeating myself thrice: the messages are confusing, the Ctrl+C handling is just utterly broken, the citation message adds to the confusion while being frankly obnoxious, and all of the delays and outputs are unnecessary in the first place as proven by literally every other program that doesn't make me wait a minute before I can use it the first time, including xargs. If Microsoft's own Windows tools did this, everybody would bash them (no pun intended) till the end of time. But since it's GNU Parallel and not Microsoft Parallel, it's Windows's fault for being slow and also my fault for having the audacity to expect better, apparently.

            • MrDOS 4 hours ago

              > the messages are confusing > while being frankly obnoxious

              Try reading some time. It's pretty great.

              > the Ctrl+C handling is just utterly broken

              That's your terminal, not parallel.

              > and all of the delays

              Your terminal is adding the delay, not parallel. parallel is just warning you about your broken terminal. You're shooting the messenger here.

    • shawn_w 13 hours ago

      >Why the hell was perl even relevant? Nothing I wrote used Perl.

      parallel is written in perl.

    • fn-mote 14 hours ago

      > Today I foolishly thought maybe I was the one who was holding it wrong […]

      Apparently goes on to describe being confused about parallel reading from the standard input?

    • tadfisher 10 hours ago

      This is pretty funny. You almost had me until the Perl part. Not going to get baited this time!

  • metadat 16 hours ago

    Every time I’ve used Gnu Parallel on a new system it required accepting a Eula, which is annoying. I stick to xargs -P99 these days and am happier.

    Utility software which has non-essential different first-run behavior is hostile to users.

  • sidewndr46 13 hours ago

    isn't parallel that tool that dumps some kind of begging message into your output each time you invoke it?

  • quotemstr 17 hours ago

    I love GNU Parallel too. I love how it's free software and how I can patch out the obnoxious citation notice.

  • jeffbee 16 hours ago

    Parallel is ok but rush is nicer and much faster.

    https://github.com/shenwei356/rush

    • foobarqux 15 hours ago

      The rush project page on Github itself says it is only "slightly" faster.

      • jeffbee 12 hours ago

        Modesty, I suppose. The larger your machine is, the more obviously faster it is.

pif 3 hours ago

> Or maybe you pipe into xargs and pray your filenames don’t have spaces

So they wrote a new tool and posted it on HN before checking the manual about null character termination... Not trustable at all, what a waste of time!

  • theideaofcoffee 3 hours ago

    Same thought I had, and that is where I stopped reading. All of these looping/enumerating techniques, including the weirdness with null-terminating strings to get around the shell's inherent treatment of spacing, are muscle memory for anyone that has spent more than 27 minutes in the shell at any point in the last 30 years. Yes it's weird because it's old. Deal with it. Why make yet another tool with yet another set of quirks?

pjmlp 9 hours ago

One thing with classical UNIX commands, is that you can expect to find them into random computers besides one's own laptop.

Not everyone has the luxury to only work with their own computer, or run random software on IT/customer managed systems.

  • ykonstant 6 hours ago

    This is why I tend to stick to POSIX shell scripting these days, avoiding even bash. The restriction does make some things too painful, though. Before the latest POSIX revision, certain workflows were comically hard to do (unless you combine sh with... M4). The 2024 revision injects some much-needed sanity, but I am wondering, will random computers have shells compliant with that revision?

    • pjmlp 5 hours ago

      Back in the day, as I mentioned around a few times, I found refuge in XEmacs, coming from the comfy Borland IDEs world.

      However, I also had to get comfortable enough with vi, because most customer systems that we would be required to access on a support visit, or remotely, only had ed and vi available.

      Even if we nowadays stick with Linux distros instead of UNIX in general, there is still the issue of what is there by default.

      Yep POSIX shell scripting can be a bit painful, which is why knowing a bit awk, perl and sed might help as well, as they tend to be available by default.

jllyhill 8 hours ago

Did you write it or did Claude code slopcoded it for you? Claude is the contributor to all your other repositories. There is a world of difference between "here's a problem that I'm really concerned with and poured all my expertise to solve it" and "I told Claude to fix it for me and now I'm gonna abandon it as soon as I'm done with the HN advertising".

ggm 18 hours ago

  find -print0 | xargs -0 -I {} "the {} iterated command"
  • somat 17 hours ago

    Xargs is fine, but openbsd has a -J option and every time I read the man page to figure out how to use it I read the -I and -J options and my brain glazes over.

    https://man.openbsd.org/xargs

    Other brain glazing obsd wierdness is it's two argument cd command, a cryptid I am unable to wrap my head around.

    https://man.openbsd.org/ksh#cd~2

    • fiddlerwoaroof 15 hours ago

      The two arg cd is pretty useful (and, zsh implements it too because cd is a shell feature): if you have parallel directory structure (e.g. the way rails does tests) you can switch from app/b/c/d to spec/b/c/f by doing `cd app spec`

    • 0xbadcafebee 17 hours ago

      I highly recommend memorizing the POSIX version of every *NIX command and sticking to them. 60% of the time, they work every time. (https://pubs.opengroup.org/onlinepubs/9799919799/idx/utiliti...)

      • JoshTriplett 15 hours ago

        I highly recommend using whatever capabilities are convenient of any utility you're running, and not caring about what other systems do unless you're actually trying to write a portable shell script.

      • fiddlerwoaroof 15 hours ago

        In a world with nix and where nearly every system has zsh, restricting yourself to the POSIX version of these tools is masochistic.

      • ggm 16 hours ago

        BSD vs linux with a standards committee inbetween.

  • koolba 18 hours ago

    And for most of those commands add a dash dash so that nothing with a dash prefix turns into options.

  • em-bee 16 hours ago

    why would you ever pipe find into xargs instead of calling -exec?

        find -exec the '{}' iterated command ';'
    • gumby 16 hours ago

      Because xargs is faster. Exec will invoke the command once per matching file (which is sometimes what you want, of course)!

      While xargs will accumulate a bunch of file names, then when it as n names will invoke the command with those names, while continuing to accumulate names until n is reached or the pipe closes.

      The size of n depends on the system, but is usually at least a thousand.

      • em-bee 16 hours ago

        the example as given does not accumulate, so that is what i worked with.

            find -exec command '{}' '+' 
        
        accumulates file arguments too. the only advantage of xargs is that you can tell it how many arguments to accumulate,̶ ̶t̶h̶e̶ ̶d̶o̶w̶n̶s̶i̶d̶e̶ ̶o̶f̶ ̶x̶a̶r̶g̶s̶ ̶i̶s̶ ̶t̶h̶a̶t̶ ̶y̶o̶u̶ ̶h̶a̶v̶e̶ ̶t̶o̶ ̶s̶p̶e̶c̶i̶f̶y̶ ̶t̶h̶e̶ ̶n̶u̶m̶b̶e̶r̶,̶ ̶w̶h̶e̶r̶e̶a̶s̶ ̶f̶i̶n̶d̶ ̶j̶u̶s̶t̶ ̶f̶i̶t̶s̶ ̶a̶s̶ ̶m̶a̶n̶y̶ ̶a̶s̶ ̶i̶t̶ ̶c̶a̶n̶.̶
        • xorcist 14 hours ago

          The problem with find is that you can't specify the maximum tolerated command line lengths, and find hasn't historically been very smart about it (it just had a compiled-in value). Apparently this is fixed in GNU find, but for older systems and other platforms that may still be an issue.

          Another feature missing in find that xargs has is the maximum processes to start at a time. find will run the commands in sequence, but in many situations you really want to run a bunch in parallel.

          • em-bee 14 hours ago

            good points that i wasn't aware of, thank you. though personally i rarely start huge operations that need to be parallelized so i prefer simplicity over speed.

        • unsnap_biceps 16 hours ago

          xargs will fit as many as it can in 128 KiB or the system limit, which is smaller, so in practice, it's almost always pretty similar default packing

      • ggm 16 hours ago

        The prior has a point because 98.89% of the time I type xargs -n 1 -0 and we are deep in the useless pipe argument which Rob Pike amongst others has rehearsed well.

        I do it because I do it, just like why I use egrep and sed in pipes along with awk.

    • t-3 16 hours ago

      Leah Neukirchen's lr and xe are very nice as a find and xargs replacement (although lr's test flag is way too complex). lr *.file | xe -s ' ... ' is a really great pattern for iterative scripting and hard to get wrong.

inigyou 3 hours ago

There's room for so much more innovation in the shell space. Microsoft did it with PowerShell - Linux seems to be stuck in the 1980s. You do have newer shells like fish, but they seem to only mess with the on-screen layout of the command prompt (while being stupid enough to still run within a terminal) and not the commands themselves.

I wonder if the generation one or two before mine grew up without shells and then had to make them and had no particular preconceptions, but my generation grew up treating the way shells work as a law of physics and thus doesn't innovate. Actually I wonder how much ossification is explained by this in general.

  • hsbauauvhabzb 3 hours ago

    Shells imo are legacy on legacy, I like bash and zsh, they’re nonsensical but they work. Powershell is undoubtedly saner, but I hate it.

    I think the whole concept could be rewritten from the ground up to be modern, but reaching critical mass and mvp is too damn difficult, and the majority of mac/Linux users probably don’t want it.

    • inigyou 3 hours ago

      If you build it and it's better, they will come. None of us here are building it.

    • anthk 3 hours ago

      It's the opposite. Plan9's rc(1) it's much saner than Bash, Ksh and PowerShell.

kfsone 9 hours ago

Careful - you start from the POSIX terminal spec, you think maybe it would be interesting having objects instead of raw text streams, and next thing you've reinvented powershell...

(*35 years living and loving sh/ksh/bash/dash etc, only tried pwsh so I could write some comparisons and slag off MS a bit; now it's my default shell on everything)

tom_alexander 4 hours ago

> Or maybe you pipe into xargs and pray your filenames don’t have spaces:

> find . -name '*.log' | xargs rm

No need for prayer:

  find . -name '*.log' -print0 | xargs --null rm
That uses null as a delimiter instead of linebreaks.
  • ninkendo 2 hours ago

    Related: I wish there were a portable way for xargs to treat newlines as delimeters but not other spaces. Since 99.9% of the time that's what you want. (GNU's has this with -d/--delimiter, but BSD's, and thus macOS's, doesn't.)

    --null/-0 is portable and helps, but it means the input has to be NUL-delimited, which is... rare. Sure, find has -print0 but (a) I wish I didn't need to do that, and (b) sometimes it's nice to just pipe `ls` output into xargs, without a `... | while read i; do echo -ne "${i}\0"; done | ...` kludge in the pipeline.

    Because spaces in filenames is just common enough to be something that will bite you eventually, but newlines in filenames are a straight evil that you can basically always say is the filename's fault.

dtj1123 29 minutes ago

Don't forget gnu parallel

mslusarz 4 hours ago

I was also sick of bash ways of dealing with lots of data, spaces in file names, etc, so I created csv-nix-tools: https://github.com/mslusarz/csv-nix-tools

E.g. removing all temporary files without the need to deal with spaces, glob limits, or silly find syntax:

csv-ls -R -c full_path . | csv-grep -c full_path -e '~$' | csv-exec -- rm -f %full_path

teo_zero 13 hours ago

Good idea but strange syntax. It would be more idiomatic if the command came first and the list of globs last.

Additionally the use of "--" is not what everybody expects: here it is used to introduce one argument, the command, while it's usually meant to introduce multiple arguments without worrying if they have a leading "-".

A possible revised syntax with command as the first argument followed by a list of globs optionally introduced by "--" would allow to enumerate all files with a leading "-", which the current syntax cannot:

  enumerate 'whatever {}' -- '-*'
I'm assuming "-f" for simplicity, but the same reasoning holds for "-L" too.
  • montag 13 hours ago

    I thought -- was typically a separator for passing a list of arguments directly to some subcommand

    • altairprime 10 hours ago

      The historical meaning is, "do not interpret command line arguments following --"; e.g. the classical form shown by `echo > -f; rm -- -f`. Git has a more complex interpretation of it and no doubt there's others, but this root interpretation remains generally sound: It acts as a boundary between 'complex and intelligent processing of @ARGV elements' and 'every remaining element of @ARGV after -- is treated a string literal without further processing'.

danlitt 3 hours ago

I think this post finally convinced me shells are awful and should be avoided.

chasil 16 hours ago

I use:

  find . -name '*.log' -print0 | xargs -0 rm
For this simple example (derived from the article), find also has a delete operator.

This null termination is now a POSIX standard.

  • ivlad 16 hours ago

    Was exactly my thought: why not to use null termination?

    Looks like a case where reading man page would have spared writing another copycat utility.

    • rmunn 13 hours ago

      And null termination is guaranteed to work, because the only two characters forbidden in Unix filenames (for most varieties of Unix, I won't guarantee there aren't some weird variants out there) are / and null.

      The only times I've needed something more than `find -print0 | xargs -0` has been when I need to apply logic to decide whether to process one of the files, in a way that's not easy to express in a `find` command. Then I write a small script with a for loop and if statements inside it.

      But more people should know about `-print0`. It's the answer to 95% of the problems with `find | xargs`.

  • Walf 14 hours ago

    Yeah, and I invariably add -r to xargs, so as not to execute the command unless results come through the pipeline.

    If one is working with whole lines of text, setting the delimiter to newline is often desirable:

    xargs -d \\n

guido26 3 hours ago

My Summary of comments: Multiple commands in use can interpret a series of string values in weird ways.

You need to establish that every program's interpretation points to the same object before processing the object through the command chain. Stopping a bad command is more important than running the good ones.

There is no tool that does this one job.

kekqqq 4 hours ago

Parallel and sed are already well-established tools.

G_o_D 12 hours ago

Atleast follow standard practice in your example

Fails : enumerate -f '.txt' -- 'wc -l {}'

For all use cases

Correct : enumerate -f '.txt' -- 'wc -l -- {}'

gfalcao 12 hours ago

> Or maybe you pipe into xargs and pray your filenames don’t have spaces…

Most of this, if not all, is fixable by adding a `export IFS=$'\n'` to your bashrc. I'm not trying to disregard your project, just point out something that took me years to learn and I currently use extensively to solve this very problem. Perhaps you didn't know about it until now... :)

  • ykonstant 6 hours ago

    Try not to do that in an uncontrolled file system. Filenames can and will have newlines, including trailing newlines, because evil people like myself inject them everywhere in our local user folders to keep sysadmins on their toes. Use the (now, finally, POSIX) -0/-print0 options for all file parsing which produces the correct behavior on all UNIX systems.

  • ivlad 9 hours ago

    Or, you use `xargs -0` for null termination instead of white space termination. `find` conveniently supports `-print0` that will use null character as separator.

  • gfalcao 11 hours ago

    Another trick that took me years to learn is to use `xargs -I` to split the results into "items"

    For example

    `ps aux | grep process-name | grep -v grep | awk '{ print $2 }' | xargs -Ieach kill -9 each`

loremm 16 hours ago

I have found that the most reliable way I like is to just construct the command externally and then pass to gnu parallel (mostly for --eta and --tmuxpane). And the great thing is, as others say, xargs -I. I prefer for shortness (and few collisoins), '@'

seq 1 10|xargs -I@ echo 'bash run.py @'|parallel -j 10

I know the echo is a little silly but then I can remove the |parallel and see if it's right. And if I don't want parallelism, I just pass to bash

vladde 8 hours ago

i often find xargs ends up biting me, and i have wanted some alternative for a while... but i don't think this is the one for me.

it feel like the syntax here is odd. it still requires me to write quote my command i want to run? unfortunately i'll have to pass on this.

personally i'd want some variant where i can still auto-complete commands and have just have {} as a placeholder. (maybe time to learn how to use xargs for real?)

scrame 9 hours ago

find -print0...|xargs -0... works for me, and i don't always want to execute something and xargs can run parallel processes. i feel like this guy never bothered to read the man page.

samtheprogram 16 hours ago

Literally just use xargs with -I {} and quotation marks?

  • laughing_man 15 hours ago

    That won't save you from weird file names, but null termination will.

    • samtheprogram 13 hours ago

      Passing "{}" handles most sane cases including spaces. If I'm doing bash that needs to be robust (rare and/or dotfiles) or know the folder/dataset has weird filenames, sure.

tom_ 5 hours ago

I never liked xargs either. My replacement is kind of like "xargs -n 1 -d '\n' -J '{}'", which experience has taught me is almost always what I want. It consumes all of its input before starting, so it knows how many files it has to process, and it can print progress to stderr and/or the terminal title as it goes. It can run each command via the shell if you want. It has --dry-run and --keep-going. It can read the file names from a file rather than stdin. I have a few more ideas for things it could do, but I haven't needed to add them yet.

(For enumerating files, I use find or dir/b/s, possibly combined with grep, then pipe the result in.)

People moaning about avoiding non-basic use of xargs and bash (and inadvertently demonstrating in many cases why some of us think that xargs sucks) miss a large part of the point, which is that it's nice to have a tool that does exactly what you want, and works in a way that's convenient for you, and isn't so widely used that you have to worry about modifying it. If you find the tool doesn't work the way you like, you can just change it. You don't have to be answerable to anybody else. I think tptacek's quite good essay could be relevant: https://sockpuppet.org/blog/2026/05/12/emacsification/

(You don't have to use LLMs for this! I wrote my program by hand, it's only like 250 lines of Python, and you could write one too. But, whichever barriers to entry prevent you from creating your own tools for yourself, using an LLM would probably lower at least some of them.)

0xbadcafebee 16 hours ago

I wanted something simpler — one consistent way to iterate over anything.

That's not actually simpler though. Simple is removing everything unnecessary. You took commands which could already do what you wanted, and added an extra program which calls them in specific ways. This will add bugs and maintenance headaches, not be portable, etc. This is added complexity.

The reason you made this script is not because you wanted simpler, you wanted easier. There's nothing wrong with that, and I'll grant you it probably is, especially for those unaccustomed to these commands. But easier != simpler. Often you'll find that simple is hard and easy is complexity deferred.

raggi 16 hours ago

in zsh you can just write:

   for x (*.sh); echo "before $x after"
or

   for x in *.sh; echo "before $x after"
or

   for x (*.sh) { echo -n "before "; echo -n $x; echo " after" }
  • PunchyHamster 16 hours ago

    missed the point. reason to use xargs or parallel are generally two: list of arguments is too long, or list of arguments that is kept in memory would take too much

    for example

        for a in `find / ` ; do echo $a ; done
    
    will take A LOT of memory, while using find's -exec, xargs, or parallel will not
    • unsnap_biceps 16 hours ago

      I looked into the source and the underlaying source uses a glob and suffers from the issue you brought up.

      https://github.com/wallach-game/bashumerate/blob/master/lib/...

    • raggi 10 hours ago

      In my experience once you get to the point where you run out of space in the glob you’re often suffering with poor performance from spawning children as well and it’s time to move to a more formal program even if it’s a script, writing out work plans and completions to list files to avoid wasted time. It’s often a great guardrail to remind you to do this

    • foobarqux 16 hours ago

      But bashenumerate doesn't do that? The parent is saying you should zsh shortloops instead of bashenumerate not zsh shortloops instead of xargs.

jeffffff an hour ago

y'all still write shell commands manually?

fhn 11 hours ago

calling it 'bashnumerate' and then have the command be 'enumerate' is confusing. find one and stick with it.

AdieuToLogic 12 hours ago

> Ever found yourself writing this?

  for f in *.txt; do
    wc -l "$f"
  done
No, because `wc` accepts multiple files. And the example given is incorrect for any file having a `.txt` suffix and whitespaces.

> Or this?

  find . -name '*.sh' -exec wc -l {} +
No.

> These all work, but each has its own syntax, its own flags, its own quirks. I wanted something simpler — one consistent way to iterate over anything.

And therein lies the proverbial xkcd standards[0] proof.

0 - https://xkcd.com/927/

caminanteblanco 18 hours ago

Obligatory XKCD: https://xkcd.com/927/

PunchyHamster 16 hours ago

I feel like it could be just alias to some particular GNU Parallel set of options

zombot 5 hours ago

> and pray your filenames don’t have spaces

`find` has `-print0` and `xargs` has the matching `-0` flag for this. So this specific argument makes me doubt the author knows their tools. Stopped reading at this point.

dboreham 13 hours ago

55 comments and nobody mentioned the incorrect English in the title?

  • fhn 11 hours ago

    I'm glad you here to save us. What would we do without you?

wallach-gameOP 19 hours ago

bashumerate — iterate over files, lines, ranges, or lists with a consistent {} syntax. No for loops, no find -exec, no xargs flags to remember. enumerate -f '*.sh' -- wc -l {} enumerate -L a b c -- 'echo {}' Under 150 lines of bash, pluable sources, NUL-safe. https://github.com/wallach-game/bashumerate

Keyboard Shortcuts

j
Next item
k
Previous item
o / Enter
Open selected item
?
Show this help
Esc
Close modal / clear selection