AI & ML interests

terminology; sentiment analysis; natural language processing; emotion detection; machine translation

Recent Activity

natalievgrafova  updated a dataset about 1 month ago
LT3/LoveHate.RU
natalievgrafova  updated a dataset about 1 month ago
LT3/LoveHate.RU
natalievgrafova  published a dataset 2 months ago
LT3/LoveHate.RU
View all activity

BramVanroy 
posted an update 10 months ago
view post
Post
782
What are currently the best multilingual models with at most 72B parameters? Are Llama 3.3 70B and Qwen 2.5 72B still king?
  • 1 reply
·
BramVanroy 
posted an update 11 months ago
view post
Post
1182
Thanks to popular request, I've just added two subsets to the CommonCrawl-Creative Commons Corpus (C5; BramVanroy/CommonCrawl-CreativeCommons) so that you do not have to do filtering manually

- C5f ( BramVanroy/CommonCrawl-CreativeCommons-fine): only retains high-quality samples that are also present in FineWeb or FineWeb-2;
- C5r (https://huggingface.co/datasets/BramVanroy/CommonCrawl-CreativeCommons-recommended): additional strict filtering that removes samples with license disagreement, non-commercial licenses, and Wikipedia samples. The latter because you should probably get those from a more reliable source that provides better parsed content.

It goes without saying that these filters lead to a massive reduction in quantity. Doc and token counts are given on the dataset pages.