Sinabis Oil – Synthetic Email Dataset
More than 1.25 million emails from the everyday business of a fictional
energy company. They span eight and a half years and four languages, and they
include real attachments and calendar invitations. Every person, company and
event is made up. The content was written by large language models, and the
demo-gen generator built the structure around it.
All figures in this file were computed from the files themselves (as of 1 Oct 2026). They are not target values from the configuration.
At a glance
| Emails | 1,251,102 (13.1 GB, .eml per RFC 5322) |
| Business cases | 376,417: 262,776 email threads and 113,641 single emails |
| Meetings | 9,470 meetings with 45,483 calendar emails, plus 3,020 standalone .ics files |
| Document attachments | 57,890: 33,467 PDF · 13,485 DOCX · 10,938 XLSX |
| Time span | January 2018 to July 2026, about 146,000 emails per year |
| Languages | English, German, French, Italian |
| People | 140 employees in 15 departments, 60 external contacts |
| Organisations | 1 main company + 39 others (group companies, suppliers, authorities, customers) |
What it is about
Sinabis Oil & Gas GmbH (Wilhelmshaven, Germany, fictional) runs drilling operations in the North Sea, a refinery, pipelines and tank farms. The emails show the day-to-day business of such an operation. There are 29 topics, including:
North Sea drilling campaign · refinery turnaround · pipeline inspection · HSE incident report · IT security incident · data protection case · audit · supplier tender · CAPEX budget · shift planning · environmental permit · offshore emergency power · HR topics and more.
About 12 % of the emails are private or informal exchanges between colleagues. Single emails are newsletters, memos, circulars, payment reminders and maintenance notices.
Who is writing?
| Group | Count | Examples |
|---|---|---|
| Sinabis Oil & Gas GmbH | 140 people, 15 departments | E&P, Drilling, Refinery, HSE, Procurement, IT, Legal/Compliance |
| Sinabis group companies | 21 | Sinabis Logistik, Sinabis Finanzdienst, Sinabis Consulting |
| Suppliers and service providers | 11 | Nordsee Rig Services, Baltic Drilling Supplies, Friesland Inspection |
| Authorities | 4 | Landesbergamt Nord, Hafenbehörde Wattkant, Umweltagentur Küstenland |
| Energy customers | 3 | Ostsee Stromwerk, Nordlicht Kraftwerke, Küstenstadt Wärme |
Activity is unevenly distributed, like in a real mailbox: about 145 internal senders write 99 % of the internal emails, some of them tens of thousands each. People only appear during their period of employment.
Language distribution
The language of each email was determined from its subject and its own text, without the quoted history, using stop-word detection. Very short texts, such as calendar replies like "Accepted: …" or "Zugesagt: …", were assigned by their subject line.
| Language | Main set | Earlier set | Total | Share |
|---|---|---|---|---|
| English | 481,765 | 199,677 | 681,442 | 54.5 % |
| German | 306,392 | 32,991 | 339,383 | 27.1 % |
| French | 102,762 | 26,694 | 129,456 | 10.3 % |
| Italian | 97,069 | 3,752 | 100,821 | 8.1 % |
| Total | 987,988 | 263,114 | 1,251,102 |
Every email is monolingual: subject, body, signature, disclaimer, legal notice and history header are all in the same language.
Attachments by language
| DOCX | XLSX | ||
|---|---|---|---|
| English | 20,021 | 6,690 | 5,323 |
| German | 7,965 | 4,084 | 3,482 |
| French | 3,365 | 1,352 | 1,128 |
| Italian | 2,116 | 1,359 | 1,005 |
Models used
In the main set, the tag at the end of the file name (…-V1.eml, …-QW.eml)
shows which model wrote the email.
| Tag | Model | Emails | Share |
|---|---|---|---|
V1–V6 |
Qwen3.6-35B-A3B (AWQ, 4-bit) | 781,765 | 62.5 % |
| (none) | DeepSeek V4 Flash (no thinking), incl. all 263,114 of the earlier set | 335,068 | 26.8 % |
QW |
Qwen 3.6 | 88,786 | 7.1 % |
| (none), calendar | no LLM, rule-based | 45,483 | 3.6 % |
The PDF, DOCX and XLSX attachments are real files. The model only supplies the content; the generator builds the files itself.
Downloads
The dataset is split into archives so you only need to download what you're interested in:
| Archive | Contents |
|---|---|
Sinabis_Demo_Dataset_emails.zip |
1,251,102 .eml files |
Sinabis_Demo_Dataset_chats.zip |
chat/messenger conversations |
Sinabis_Demo_Dataset_appointments.zip |
standalone .ics calendar files |
Each archive keeps the internal folder structure shown below. README.md and
docs/ are plain files in the repo, not zipped.
Folder structure
.
├── README.md
├── docs/
│ └── generation/
│ ├── prompts.json
│ ├── chat_prompts.json
│ └── company.json
├── Sinabis_Demo_Dataset_emails.zip
│ └── batch-0001 … batch-0031/
├── Sinabis_Demo_Dataset_chats.zip
│ └── batch-0001 … batch-0002/
└── Sinabis_Demo_Dataset_appointments.zip
The batch folders only split the data so that no single folder holds a million files. They carry no meaning.
How everything is linked
Email threads
<thread-id>-<NN>-<TAG>.eml
│ │ └─ model (see above), missing for DeepSeek and calendar
│ └────── position in the thread (01, 02, …)
└────────────────── SHA-256 over the whole thread, same for all its emails
- All emails of one case share the same file name prefix.
- The headers link the emails:
Message-ID,In-Reply-ToandReferenceseach point to the previous email. - Every email contains the quoted history (
> …) of all earlier emails, so a single file already shows the whole thread. - 99 % of threads have 3 to 6 emails.
Attachments
- Attachments are embedded in the email (MIME, Base64), not stored as separate files next to it.
- The document is the source the email refers to, for example an inspection report or a budget sheet, not a copy of the email text. It has the same language and the same date as the email.
Meetings
- A meeting consists of several emails with an
.icsattachment: the invitation (REQUEST) to all attendees, acceptances or declines (REPLY) to the organiser, and possibly a cancellation (CANCEL). - These emails share the same file name prefix. 9,470 meetings produce 45,483 calendar emails this way.
- The standalone
.icsfiles inSinabis_Demo_Dataset_appointments.zipare meetings without an email (appt_<no>_<date>_<time>.ics).
How timestamps are created
Timestamps are set by the generator, not by the language model. Otherwise models tend to date almost everything on the same day.
- Start time: For each case the generator picks a random day between 1 Jan 2018 and 31 Jul 2026. The time falls within office hours, peaking at 10 am; only 4 % of cases fall on a weekend.
- Thread: The model writes the replies with plausible gaps between them. The whole thread is then shifted onto the start time. Replies always come after the previous email, and gaps of more than 30 days are capped.
- Dates in the text: Dates, months and years mentioned in the email text are shifted as well, so text and header stay consistent.
- People: Every person has an employment or business-relationship window. Someone hired in 2024 does not send emails in 2019.
- Attachments carry the date of their email.
The Date header is in UTC (+0000).
Emails per year
| 2018 | 2019 | 2020 | 2021 | 2022 | 2023 | 2024 | 2025 | 2026 (to July) |
|---|---|---|---|---|---|---|---|---|
| 148,256 | 146,377 | 143,558 | 146,887 | 143,227 | 146,155 | 146,034 | 145,691 | 84,790 |
Notes
- Fictional: All names, companies, addresses, phone numbers and events are made up. A brand filter removes real company names from the texts.
- Earlier set: It is more heavily English (76 %), has hardly any DOCX/XLSX and no calendar emails. Linking via file names and headers works the same way.
- Names without addresses: In about 4 % of the emails,
FromorTocontain only names without an email address, because the model answered that way. - Address variants: About 1 % of internal emails come from addresses
outside the staff list (e.g.
koenig.farid@…). These are spelling variants produced by the model, much like typos in real mailboxes. - Edge dates: 127 emails are dated late 2017.
Methodology
docs/generation/ contains the prompt templates and configuration used to
prompt the LLMs, published for transparency and reproducibility.
License
This dataset is released under CC BY 4.0. You may copy, modify, distribute and use it for any purpose, including commercially, as long as you give appropriate credit, for example:
Sinabis GmbH (2026). Sinabis Oil – Synthetic Email Dataset. https://huggingface.co/datasets/sinabis-group/sinabis-synthetic-corporate-dataset
The email content was generated using Qwen 3.6 (open weights) and DeepSeek V4 Flash (via API). Generated text is treated as free of third-party rights; the underlying models' own terms of use apply to the act of generation, not to redistribution of this dataset.
Citation
If you use this dataset, please cite it as:
Sinabis GmbH (2026). Sinabis Oil – Synthetic Email Dataset.
https://huggingface.co/datasets/sinabis-group/sinabis-synthetic-corporate-dataset
BibTeX:
@misc{sinabis_email_dataset_2026,
author = {{Sinabis GmbH}},
title = {Sinabis Oil -- Synthetic Email Dataset},
year = {2026},
url = {https://huggingface.co/datasets/sinabis-group/sinabis-synthetic-corporate-dataset}
}
- Downloads last month
- 90