sinabis-group/sinabis-synthetic-corporate-dataset · Datasets at Hugging Face

7 min read Original article ↗

Sinabis Oil – Synthetic Email Dataset

More than 1.25 million emails from the everyday business of a fictional energy company. They span eight and a half years and four languages, and they include real attachments and calendar invitations. Every person, company and event is made up. The content was written by large language models, and the demo-gen generator built the structure around it.

All figures in this file were computed from the files themselves (as of 1 Oct 2026). They are not target values from the configuration.

At a glance

Emails 1,251,102 (13.1 GB, .eml per RFC 5322)
Business cases 376,417: 262,776 email threads and 113,641 single emails
Meetings 9,470 meetings with 45,483 calendar emails, plus 3,020 standalone .ics files
Document attachments 57,890: 33,467 PDF · 13,485 DOCX · 10,938 XLSX
Time span January 2018 to July 2026, about 146,000 emails per year
Languages English, German, French, Italian
People 140 employees in 15 departments, 60 external contacts
Organisations 1 main company + 39 others (group companies, suppliers, authorities, customers)

What it is about

Sinabis Oil & Gas GmbH (Wilhelmshaven, Germany, fictional) runs drilling operations in the North Sea, a refinery, pipelines and tank farms. The emails show the day-to-day business of such an operation. There are 29 topics, including:

North Sea drilling campaign · refinery turnaround · pipeline inspection · HSE incident report · IT security incident · data protection case · audit · supplier tender · CAPEX budget · shift planning · environmental permit · offshore emergency power · HR topics and more.

About 12 % of the emails are private or informal exchanges between colleagues. Single emails are newsletters, memos, circulars, payment reminders and maintenance notices.

Who is writing?

Group Count Examples
Sinabis Oil & Gas GmbH 140 people, 15 departments E&P, Drilling, Refinery, HSE, Procurement, IT, Legal/Compliance
Sinabis group companies 21 Sinabis Logistik, Sinabis Finanzdienst, Sinabis Consulting
Suppliers and service providers 11 Nordsee Rig Services, Baltic Drilling Supplies, Friesland Inspection
Authorities 4 Landesbergamt Nord, Hafenbehörde Wattkant, Umweltagentur Küstenland
Energy customers 3 Ostsee Stromwerk, Nordlicht Kraftwerke, Küstenstadt Wärme

Activity is unevenly distributed, like in a real mailbox: about 145 internal senders write 99 % of the internal emails, some of them tens of thousands each. People only appear during their period of employment.

Language distribution

The language of each email was determined from its subject and its own text, without the quoted history, using stop-word detection. Very short texts, such as calendar replies like "Accepted: …" or "Zugesagt: …", were assigned by their subject line.

Language Main set Earlier set Total Share
English 481,765 199,677 681,442 54.5 %
German 306,392 32,991 339,383 27.1 %
French 102,762 26,694 129,456 10.3 %
Italian 97,069 3,752 100,821 8.1 %
Total 987,988 263,114 1,251,102

Every email is monolingual: subject, body, signature, disclaimer, legal notice and history header are all in the same language.

Attachments by language

PDF DOCX XLSX
English 20,021 6,690 5,323
German 7,965 4,084 3,482
French 3,365 1,352 1,128
Italian 2,116 1,359 1,005

Models used

In the main set, the tag at the end of the file name (…-V1.eml, …-QW.eml) shows which model wrote the email.

Tag Model Emails Share
V1–V6 Qwen3.6-35B-A3B (AWQ, 4-bit) 781,765 62.5 %
(none) DeepSeek V4 Flash (no thinking), incl. all 263,114 of the earlier set 335,068 26.8 %
QW Qwen 3.6 88,786 7.1 %
(none), calendar no LLM, rule-based 45,483 3.6 %

The PDF, DOCX and XLSX attachments are real files. The model only supplies the content; the generator builds the files itself.

Downloads

The dataset is split into archives so you only need to download what you're interested in:

Archive Contents
Sinabis_Demo_Dataset_emails.zip 1,251,102 .eml files
Sinabis_Demo_Dataset_chats.zip chat/messenger conversations
Sinabis_Demo_Dataset_appointments.zip standalone .ics calendar files

Each archive keeps the internal folder structure shown below. README.md and docs/ are plain files in the repo, not zipped.

Folder structure

.
├── README.md
├── docs/
│   └── generation/
│       ├── prompts.json
│       ├── chat_prompts.json
│       └── company.json
├── Sinabis_Demo_Dataset_emails.zip
│   └── batch-0001 … batch-0031/
├── Sinabis_Demo_Dataset_chats.zip
│   └── batch-0001 … batch-0002/
└── Sinabis_Demo_Dataset_appointments.zip

The batch folders only split the data so that no single folder holds a million files. They carry no meaning.

How everything is linked

Email threads

<thread-id>-<NN>-<TAG>.eml
 │           │    └─ model (see above), missing for DeepSeek and calendar
 │           └────── position in the thread (01, 02, …)
 └────────────────── SHA-256 over the whole thread, same for all its emails
  • All emails of one case share the same file name prefix.
  • The headers link the emails: Message-ID, In-Reply-To and References each point to the previous email.
  • Every email contains the quoted history (> …) of all earlier emails, so a single file already shows the whole thread.
  • 99 % of threads have 3 to 6 emails.

Attachments

  • Attachments are embedded in the email (MIME, Base64), not stored as separate files next to it.
  • The document is the source the email refers to, for example an inspection report or a budget sheet, not a copy of the email text. It has the same language and the same date as the email.

Meetings

  • A meeting consists of several emails with an .ics attachment: the invitation (REQUEST) to all attendees, acceptances or declines (REPLY) to the organiser, and possibly a cancellation (CANCEL).
  • These emails share the same file name prefix. 9,470 meetings produce 45,483 calendar emails this way.
  • The standalone .ics files in Sinabis_Demo_Dataset_appointments.zip are meetings without an email (appt_<no>_<date>_<time>.ics).

How timestamps are created

Timestamps are set by the generator, not by the language model. Otherwise models tend to date almost everything on the same day.

  1. Start time: For each case the generator picks a random day between 1 Jan 2018 and 31 Jul 2026. The time falls within office hours, peaking at 10 am; only 4 % of cases fall on a weekend.
  2. Thread: The model writes the replies with plausible gaps between them. The whole thread is then shifted onto the start time. Replies always come after the previous email, and gaps of more than 30 days are capped.
  3. Dates in the text: Dates, months and years mentioned in the email text are shifted as well, so text and header stay consistent.
  4. People: Every person has an employment or business-relationship window. Someone hired in 2024 does not send emails in 2019.
  5. Attachments carry the date of their email.

The Date header is in UTC (+0000).

Emails per year

2018 2019 2020 2021 2022 2023 2024 2025 2026 (to July)
148,256 146,377 143,558 146,887 143,227 146,155 146,034 145,691 84,790

Notes

  • Fictional: All names, companies, addresses, phone numbers and events are made up. A brand filter removes real company names from the texts.
  • Earlier set: It is more heavily English (76 %), has hardly any DOCX/XLSX and no calendar emails. Linking via file names and headers works the same way.
  • Names without addresses: In about 4 % of the emails, From or To contain only names without an email address, because the model answered that way.
  • Address variants: About 1 % of internal emails come from addresses outside the staff list (e.g. koenig.farid@…). These are spelling variants produced by the model, much like typos in real mailboxes.
  • Edge dates: 127 emails are dated late 2017.

Methodology

docs/generation/ contains the prompt templates and configuration used to prompt the LLMs, published for transparency and reproducibility.

License

This dataset is released under CC BY 4.0. You may copy, modify, distribute and use it for any purpose, including commercially, as long as you give appropriate credit, for example:

Sinabis GmbH (2026). Sinabis Oil – Synthetic Email Dataset. https://huggingface.co/datasets/sinabis-group/sinabis-synthetic-corporate-dataset

The email content was generated using Qwen 3.6 (open weights) and DeepSeek V4 Flash (via API). Generated text is treated as free of third-party rights; the underlying models' own terms of use apply to the act of generation, not to redistribution of this dataset.

Citation

If you use this dataset, please cite it as:

Sinabis GmbH (2026). Sinabis Oil – Synthetic Email Dataset.
https://huggingface.co/datasets/sinabis-group/sinabis-synthetic-corporate-dataset

BibTeX:

@misc{sinabis_email_dataset_2026,
  author = {{Sinabis GmbH}},
  title  = {Sinabis Oil -- Synthetic Email Dataset},
  year   = {2026},
  url    = {https://huggingface.co/datasets/sinabis-group/sinabis-synthetic-corporate-dataset}
}
Downloads last month
90