Instructions to use dronefreak/bdd100k-scenario-efficientvit_b3 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- timm
How to use dronefreak/bdd100k-scenario-efficientvit_b3 with timm:
import timm model = timm.create_model("hf-hub:dronefreak/bdd100k-scenario-efficientvit_b3", pretrained=True) - Notebooks
- Google Colab
- Kaggle
EfficientViT-B3 Finetuned on BDD100K Scenario Classification
Fine-tuned EfficientViT-B3 image classifier on the BDD100K Scenario Classification dataset, trained and evaluated as part of BDD100K-Toolkit, a dependency-clean toolkit for preparing BDD100K, training models on it and evaluating them with the same metrics on the same splits.
7-class driving-scenario classification (city street / highway / residential / parking lot / gas stations / tunnel / unknown) derived from BDD100K's per-image attributes.scene field. Unofficial task; follows the Kaggle dataset of the same name.
Usage
import timm, torch
from huggingface_hub import hf_hub_download
from PIL import Image
from torchvision import transforms as T
ckpt = torch.load(
hf_hub_download("dronefreak/bdd100k-scenario-efficientvit_b3", "best.pt"), map_location="cpu", weights_only=True
)
model = timm.create_model(
ckpt["model_name"], pretrained=False, num_classes=len(ckpt["class_names"])
)
model.load_state_dict(ckpt["state_dict"])
model.eval()
prep = T.Compose(
[T.Resize((ckpt["imgsz"],) * 2), T.ToTensor(), T.Normalize(ckpt["mean"], ckpt["std"])]
)
with torch.no_grad():
probs = model(prep(Image.open("street.jpg").convert("RGB"))[None]).softmax(1)[0]
print(ckpt["class_names"][probs.argmax()], f"{probs.max():.1%}")
Results
Evaluated on the test split (10000 images).
| Metric | Value |
|---|---|
| Macro F1 | 57.70% |
| Balanced accuracy | 53.59% |
| Macro precision | 71.54% |
| Macro recall | 53.59% |
| Top-1 accuracy | 79.32% |
| Top-5 accuracy | 99.97% |
Per class
| Class | Precision | Recall | F1 | Test images |
|---|---|---|---|---|
| city street | 81.10% | 88.30% | 84.55% | 6112 |
| gas stations | 66.67% | 28.57% | 40.00% | 7 (few) |
| highway | 79.15% | 70.63% | 74.65% | 2499 |
| parking lot | 70.00% | 42.86% | 53.16% | 49 (few) |
| residential | 68.66% | 57.70% | 62.71% | 1253 |
| tunnel | 85.19% | 85.19% | 85.19% | 27 (few) |
| unknown | 50.00% | 1.89% | 3.64% | 53 |
Model Zoo
All runs below were evaluated on the same test split, sorted by top-1 accuracy.
| Model | Top-1 | Macro F1 | Balanced acc | Macro precision |
|---|---|---|---|---|
| EfficientViT-B3 | 79.32% | 57.70% | 53.59% | 71.54% |
| YOLO11n | 78.58% | 49.47% | 46.05% | 60.98% |
| YOLO11s | 78.58% | 49.18% | 45.24% | 61.98% |
| MobileNetV4-Conv-Small | 78.20% | 52.89% | 48.69% | 61.41% |
| MobileNetV4-Conv-Large | 78.15% | 50.00% | 45.48% | 63.03% |
| EfficientViT-B1 | 78.01% | 54.15% | 49.41% | 70.13% |
| TinyViT-21M | 77.99% | 62.52% | 66.21% | 59.66% |
| YOLO26s | 77.81% | 52.24% | 49.23% | 58.28% |
| ConvNeXt-Atto | 77.44% | 61.06% | 60.34% | 67.92% |
| YOLOv8s | 77.15% | 48.43% | 46.64% | 51.16% |
| ResNet-18 | 77.14% | 47.29% | 44.83% | 54.25% |
| YOLO26n | 76.78% | 48.33% | 43.94% | 58.99% |
| ConvNeXt-Tiny | 76.69% | 56.88% | 59.82% | 54.60% |
| YOLOv8n | 76.13% | 46.40% | 43.26% | 54.14% |
| EfficientViT-B0 | 75.94% | 46.10% | 42.64% | 53.74% |
Training
| Setting | Value |
|---|---|
| Epochs (max) | 50 |
| Epochs (trained) | 21 |
Best epoch (best.pt) |
11 |
best.pt chosen by |
macro_f1 on the valid split |
| Early stopping patience | 10 |
| Batch size | 32 |
| Image size | 224 |
| Optimizer | auto, resolved to AdamW (peak lr 3e-04) |
| Weights | EMA |
Dataset
dronefreak/BDD100K-Scenario-Classification holds the prepared splits these models were trained and evaluated on.
Limitations
- Unofficial task: labels are BDD100K's per-image attributes, not a benchmark with a public leaderboard, so scores are only comparable with other models evaluated on this split.
- Not the official test set:
testhere is BDD100K's official validation split (the official test labels are not released) andvalidis a seeded 15% slice of the official train split. - Heavily imbalanced: rare classes have very few test images, so their per-class scores are noisy. Prefer macro F1 over accuracy.
- Images are US dashcam frames; generalization to other regions or camera setups is untested.
- BDD100K is released for non-commercial research and education. Check its license before any use of these weights beyond research.
sceneis a coarse per-image label; frames showing several settings (e.g. a highway entering a tunnel) get a single class.
License
The weights are released under the license in the metadata above. They were trained on BDD100K, which is free for non-commercial research and education; commercial use needs separate permission (see https://www.bdd100k.com/). The prepared dataset on the Hub is tagged license: other, and the original BDD100K terms still apply.
Citation
Model
@article{cai2022efficientvit,
title={EfficientViT: Multi-Scale Linear Attention for High-Resolution Dense Prediction},
author={Cai, Han and Li, Junyan and Hu, Muyan and Gan, Chuang and Han, Song},
journal={arXiv preprint arXiv:2205.14756},
year={2022}
}
Dataset
@article{yu2018bdd100k,
title={BDD100K: A Diverse Driving Dataset for Heterogeneous Multitask Learning},
author={Yu, Fisher and Chen, Haofeng and Wang, Xin and Xian, Wenqi and Chen, Yingying and Liu, Fangchen and Madhavan, Vashisht and Darrell, Trevor},
journal={arXiv preprint arXiv:1805.04687},
year={2018}
}
Files
best.ptmetrics.jsonresults.csvassets/demo_banner.mp4assets/demo_banner_poster.jpgREADME.md
Reproduce
Trained and evaluated with BDD100K-Toolkit: bdd100k-evaluate --dataset bdd100k-scenario --checkpoint <weights> --data-dir <prepared dir> --split test.
- Downloads last month
- -
Model tree for dronefreak/bdd100k-scenario-efficientvit_b3
Base model
timm/efficientvit_b3.r224_in1kDataset used to train dronefreak/bdd100k-scenario-efficientvit_b3
Collections including dronefreak/bdd100k-scenario-efficientvit_b3
Papers for dronefreak/bdd100k-scenario-efficientvit_b3
EfficientViT: Lightweight Multi-Scale Attention for On-Device Semantic Segmentation
BDD100K: A Diverse Driving Dataset for Heterogeneous Multitask Learning
Evaluation results
- Top-1 accuracy (test split) on BDD100K Scenario ClassificationBDD100K-Toolkit79.320
- Macro F1 (test split) on BDD100K Scenario ClassificationBDD100K-Toolkit57.700
- Macro precision (test split) on BDD100K Scenario ClassificationBDD100K-Toolkit71.540
- Macro recall (test split) on BDD100K Scenario ClassificationBDD100K-Toolkit53.590