Settings

Theme

Reverse Engineering Unknown File Formats with ImHex

werwolv.net

214 points by carlos-menezes · 41 comments

Reader

17 threads
kstrauser

Story time: I worked the night shift at a motel during college. "Bored to tears" was an understatement. I was poking around the front desk computer system one night, as one does, and found its data file. Lacking anything else to do, I wrote a little hex dumper in BASIC so I could explore the file. The first thing I noted was that customer names were spaced exactly N bytes apart. Oooh, fixed records! Then I spent the next week or so reverse engineering the DB file format. I'd twiddle a field in a random customer record then look at their record in the hex dump to see what changed, then update my notes.

Eventually I wrote another little BASIC program to run analytics, like which customers spent the most with us, and auto-fix mistyped names ("This says 'Bb Jones'. Did you mean 'Bob Jones'?") by writing directly to the file.

I got a pretty sweet little bonus for my hackery.

  • threecheese

    I did something just like this at my first post-college job in the 90s. I found some weird backup files on the NT LAN, had similar observations as you, but wrote my extractor in C (casting records into structs - my first C program) using a “bcc” floppy disk. Turned out to be the entire insurance company Btrieve database, which I could dump into csv to load into MS Access.

    This “database” allowed me to automate nearly my entire job - finding and exercising test cases for policy bug reports with a calculator - which gave me time to experiment with software development. I’d gone to college to be a writer, but this was waaaay cooler.

    I didn’t get a bonus - heck I’d have been fired for sure - but that time spent learning made my entire career possible. Hoping to get another ten years before AI eats it all.

  • noduerme

    All was fun... sorry (about two decades younger) when I was a junior web designer in the late 90s, keyword designer, not coder, I got approached with a freelance job to build an online store, which I knew nothing about doing. I knew some PHP, mainly how to write stuff in and out of files and inline variables into html. Naturally, I said yes to the job even though I was totally unqualified. I mean, who was qualified? I had to build a whole login portal from scratch for them to upload products, and images of the products, and sub-sub-products, and add-ons and prices and descriptions, and then have all of that reflect on the public website where people would buy them by adding them to a shopping cart which was basically a very long GET query that kept getting added to.

    So but words like "database" and "sql" scared the shit out of me at the time. I was doing most of my coding out of a starbucks in Eagle Rock and there was one other guy there on a laptop, John P, who kept looking at what I was writing and telling me I had to learn databases.

    What I ended up coding was basically a database based on text files and pages and a custom encoding methods and a system for locking. Imagine a year later in 1999 when I realized this had all been solved and I'd been too stubborn to try it or understand what it did...

    Fast forward a few years and someone put two sequential asterisks in in the name of a product and blew up the site and my encoding system, since they'd accidentally stumbled on the row break symbol. lol. Good times.

  • inventor7777

    How did your program suggest auto-fixed names? A manually entered list of common names, or did it go by previous records?

    BASIC is over 40 years older than me, and I have never played with it, thus my question :-)

    • kstrauser

      BASIC, while not a language I'd use for any reason today, is Turing complete. You can write any program in BASIC that you could write in Rust. Of course, you'd mainly do that for the same reason you'd write something in, say, Brainfuck: mainly to see if you can. But when that's all we had, we could get pretty creative with it.

      It went by previous records. I'd read something about Levenshtein distance somewhere, and implemented my own half-assed version of it that was surely slower, more complicated, and less accurate. It still got the job done, though. I showed it to a coworker, who described it as "magic", and I floated on that compliment for a long time after.

      I was pretty proud of it, although the resulting QBasic code was pretty awful in retrospect. It was probably my first commercial-scale project, built with the software that happened to ship with the computer.

truekonrads

My few rules are: 1) What's the entropy? This helps with encrypted/compressed. You probably need to overcome this? 2) What's the context? Apps are developed in context and context suggests formats. Is it C? Then expect structs. Does it need to go over the wire? Expect run-length encodings. Python/JS - JSON/pickle. 3) There are broadly speaking only a few ways you can read data back into memory: fixed layout, run-length encoding, terminators. A great case-study is ASN.1, once you've built an ASN.1 parser by hand no file format will ever daunt you.

issung

I finally dipped into using ImHex when doing some reversing of the PS2 memory card format for https://ps2iodb.com.

The program is so good, and has enough of a reputation that I knew to reach for it even though I know nothing about reversing, and that was 3-4 years ago now! Also shoutout to the awesome GUI library is uses; https://github.com/ocornut/imgui

Noxwizard

I'm always on the lookout for a good hex editor. I've tried ImHex a few times and it doesn't quite fit the need I have and I prefer HxD for the simple things. The Pattern Language is pretty neat, I hadn't looked into that before. For the complicated things, it's not always a linear process like in the case of the blog article. Mainly, there's often lots of gaps in my structure until I figure them out. Some tools like ImHex or Kaitai, you can put in placeholder fields between two known objects, but that's kind of annoying. 010 Editor lets you highlight and put colored bookmarks, which kind of works. The workflow I've ended up using for reversing in-memory data structures is to get a hexdump of the memory, take a screenshot and then put it in OneNote. Then I can draw colored boxes or highlight and put annotations off on the side with what things are or notes. If a field is a memory address to another related and unknown structure, I often will then get a short dump of that memory and put a screenshot of it on the same page with an arrow to it. It's pretty kludgy, but I like the freeform nature of it. It's basically a RE notebook, just without any useful RE built-ins. If anyone has ideas or recommendations on better discovery workflows, I'd love to hear them.

vivzkestrel

- stupid question: do you have any ideas how to go about doing this on a .unr file

- it comes from a game called splinter cell conviction made in custom unreal 2.5

- I have neither been able to get UE-Explorer or UEViewer to work with it

- I can send you the file, I just want to know how to change enemy AI spawn types on it

  • NegativeLatency

    Try an LLM, either to do it directly or to guide your own explorations

    • vivzkestrel

      as a guy who doesnt have experience fiddling with this kinda stuff, what is the step by step process if any to go about figuring out how an unknown file format is actually made of and modifying it?

      - the post itself quotes "I usually couldn’t really give them a good answer except, “Look at the decompiled code of whatever program reads/writes these files and work backwards from there.”

      - i dont have the program and that is the real challenge

      • subscribed

        I would start with something like this: "Today's session will be a bit different and very exciting to me.

        I have this game, Game Name exact version, which I love and I logged hundreds of hours in it, however there's this little detail that bothers me like a pebble in a shoe: (details of the problem here). I think that it's all encoded in the file (file name), but it's binary and I have no idea how to access it :(

        Could you perhaps guide me how to analyse and "decode" this file so I could later change this behaviour? I don't have any experience with reverse engineering but I'm excited to try.

        All game files are in /full/path/to/the/game

        Thanks!"

        Works for me.

        Be excited, frame it as a learning opportunity, make sure it's obvious its for your own enjoyment (not for distribution).

  • mschuster91

    Might want to look into FModel (https://fmodel.app/) and how the Star Wars Republic Commando modders do things [1][2].

    [1] https://www.moddb.com/games/star-wars-republic-commando/tuto...

    [2] https://github.com/SWRC-Modding/CT

yapfrog

I tried to do a similar thing last year to datamine a gacha game's assets and somehow did it without knowing the existence of tools like IDA/Ghidra. They're using a custom db format that was XORed with a key. Iirc it was a bizarre format, it's actually an archive where each column is a file containing all the rows of that column. So a table is just a folder full of these files. Maybe it's a common thing idk. I got a partially working XOR key so I used ImHex to look at the partially decrypted db file and manually looked at the decrypted hex values to "correct" them.

I wish I read something like this article at that time!

jonhohle

I’ve had a lot of luck with ImHex looking at various binary formats and even using it to aid in writing some file magic.

It would be really nice to be able to roundtrip C headers with it. The syntax is close, but often I older formats there is a 1:1 mapping with C structs that doesn’t quite match the imhex syntax.

hmartin

Spoiler: half the time, that unknown file format is a zip containing a sqlite db

  • userbinator

    The older the system/the more changes it's had, especially if it's "enterprise software", the more insane things tend to be; in one memorable case, it was a .zip containing JSON, which then contained base64-encoded XML, wrapping another base64-encoded .zip containing more XML, and in that XML was base16-encoded ASN.1 PER with the actual payload (encoded in a proprietary format). I won't disclose specifics, but perhaps someone may recognise that from its description; suffice to say the industry that was related to started using computers in the days of mainframes and punched cards.

  • orphea

    Uhm... a (g)zip - yes, but a sqlitedb inside - no, not that often, in my experience.

  • butvacuum

    for games its more often .net BinarySerializer.

xiej

I've been reverse engineering some old Minecraft Legacy Console Edition save files with ImHex and the pattern templates work pretty well, though debugging is sometimes painful (e.g. reading runtime-sized arrays can just overflow the file and error without showing previously decoded structs). Great tool though. LLMs are, of course, scarily good at reverse engineering and spotting patterns that take me way too much effort to spot.

Lucasoato

I’m going to try that on the newest outlook binary format used to store the emails locally... why couldn’t they stick to the plain simple sqlite?

polishdude20

I've always wondered how something like an Autodesk Inventor sldprt file was formatted so that I could write a program to manipulate those files.

I assumed it was proprietary and left it at that.

A few years later, I saw that Onshape allows importing sldprt files into their system and I wonder how they were able to provide that ability without some proprietary knowledge transferred

shitcoder

Love imhex and recognised the author straight away. Used it plenty of times got save file reverse engineering to develop save editors

max979

ImHex makes tackling proprietary formats so much less painful. Seriously a lifesaver for quickly mapping out structures.

cdbattags

protobuf derivatives must be littered throughout

tbt to reverse engineering the original PokemonGo rpc

thataccount

If only it were BlackIce, it would be perfect.

throneitlol

One trick that saves a lot of time before you write any pattern: generate two saves that differ in exactly one known way (buy one item, then diff), and byte-diff them. The changed offsets are usually the field you care about, and the size of the changed run tells you the width. Doing this a few times gives you a rough field map without guessing at structure.

The other thing worth checking early is whether the file is compressed or checksummed. If entropy is flat and high across the whole file it is probably deflate or an encrypted blob, and hex staring will get you nowhere until you unwrap it. And if there is a 4 byte value near the header or footer that changes on every save even when the data is identical, treat it as a CRC or timestamp before assuming it is data. Writing back an edited file that fails the checksum is the classic reason a patched save silently refuses to load.

someperson

Damn a video game spoiler warning, guess I can't read the rest of the article.

I may feed it into an LLM to strip out the spoilers actually

Keyboard Shortcuts

j
Next item
k
Previous item
o / Enter
Open selected item
?
Show this help
Esc
Close modal / clear selection