I released my own benchmark sets behind emotional-memory, my library for mood-aware recall in LLM agents. The card opens with the limits: the main set is built to favour the method, and the negative results are listed up front, e.g. on LoCoMo AFT scores F1 0.168 against 0.271 for a naive RAG baseline. Use it to check where affect-conditioned retrieval helps and where it doesn't. gianlucamazza/emotional-memory-benchmarks
- an hourly news desk with exhausted overnight anchors - a game show with a reigning champion - small-claims court and a cursed cooking show - a talk show hosted by a houseplant or a microwave - a 1 a.m. call-in hotline - an β80s sitcom on a decaying VHS tape - 4 a.m. infomercials that slide into existential dread - wordless generative music videos - many more shows and content added daily...
Fully AI-run: AI writes the scripts, invents new characters every episode, designs each one a unique voice, directs the camera, cuts the edit, schedules the day and runs the broadcast.
Nothing ever airs twice: a memory of everything thatβs aired blocks repeated premises, jokes and characters, so every minute is new.
Real facts, delivered weird: the news, trivia and explainers are sourced and cited; the comedy is the chaos around them.
Its own engine: a custom pixel renderer with a virtual camera, lip-synced close-ups and 50+ animated sets, all running in a browser on one cloud box.
Itβs weird, itβs for adults, and it never stops. Happy to answer questions about how it works.
Most NSFW classifiers break the second an image touches the internet.
They look great on pristine benchmarks, but in the wild, every social platform aggressively recompresses, downsamples, and degrades images. The moment JPEG or WebP compression artifacts show up, confidence collapses and false positives spike.
SafeScan was built to survive actual platform pipelines. Trained on 34,000 images under almost every major social media compression profile using a Vision Transformer backbone (google/vit-base-patch16-224). Instead of blunt binary filtering, it breaks decisions down across 5 clear categories:
β’ safe β’ drawing β’ sexy β’ hentai β’ porn
The result is a moderation model that actually generalizes to real-world internet feeds instead of fragile, uncompressed datasets. Open-weight and available on Hugging Face:
The best result so far is 8 of the 14 playable levels cleared in one continuous run. Two models have cleared eight levels so far: - Qwen/Qwen3.8-Flash-Next - Cloudflare/clef
If you'd like me to try a specific model, please name it in the comments.
You can also try beating the game with a model of your choice. The project with instructions for running the challenge with different models/engines is here: https://github.com/felladrin/ai-plays-cat-goric (PRs welcome!)
And attached is a 2-minute recording of Clef clearing eight levels (the game pauses while the model thinks, so the video is time-warped).
This demo showcases the OneDecision-VisionGuard family of multimodal image classification models for detecting NSFW and other sensitive visual content, with structured JSON reasoning, improved accuracy, and better handling of edge cases such as sensitive imagery, uncensored analysis, scene descriptions, and classification reasoning.
Most deepfake audio detectors are quietly cheating.
They donβt really listen to the speech β they just look at how long the embedding vector is. Once they figure that out, accuracy looks great on paper and falls apart in the wild.
AIRealNet-Audio was built to stop that shortcut. It forces every feature onto the unit hypersphere (twice) so the model can only use direction, not magnitude. Trained on speech from 100+ different TTS and voice-cloning systems, plus real human recordings under heavy compression and noise.
The result is a detector that actually has to learn the artifacts instead of gaming the feature space.