Together AI (@togethercompute) on X

X (formerly Twitter) ·

2 min read Original article ↗

Post

Post

Together AI on X: "We are excited to release RedPajama-Data-v2: 30 trillion filtered & de-duplicated tokens from 84 CommonCrawl dumps, 25x larger than our first dataset. It exposes a diverse range of quality annotations so you can slice & weight the data for LLM training. https://t.co/5zT0C3QBjQ"

  • user avatar

    We are excited to release RedPajama-Data-v2: 30 trillion filtered & de-duplicated tokens from 84 CommonCrawl dumps, 25x larger than our first dataset. It exposes a diverse range of quality annotations so you can slice & weight the data for LLM training. together.ai/blog/redpajama…

  • user avatar

    Since releasing RedPajama-v1 in March, we have been thrilled by the traction, with over 500 models created on the data. We've received incredible feedback from the open-source AI community and worked to address it in this release!

    user avatar

    The dataset covers 5 languages, with 40+ pre-computed data quality annotations that can be used for further filtering and weighting. Here is one example of how to filter RedPajama-Data-v2 in a similar way as Gopher:

    user avatar

    user avatar

    All data processing scripts are open source and available on

    @github

    .

    user avatar

    We are appreciative to so many partners and collaborators that together are pushing forward the frontier of open LLM models. Thank you to the OLMo team at

    @allen_ai

    and friends at

    @OpenGPTX

    for the insightful discussions about datasets and data quality!

  • user avatar

    Is it somehow textified or you keep all the html, js, css etc.? Specifically people have been asking about json-ld schema markup