Microsoft's DeBERTa-v3-small Beats RoBERTa Using Half the Parameters
Takeaways − DeBERTa-v3-small is a 44M-parameter English encoder with a 128K vocabulary from Microsoft Research. Uses ELECTRA-style replaced-token-detection training instead of masked language modeling for better sample efficiency.

- Takeaways − DeBERTa-v3-small is a 44M-parameter English encoder with a 128K vocabulary from Microsoft Research.
- Scores 88.3 on MNLI-m and 82.8 F1 on SQuAD 2.0, matching models roughly twice its size.
- Why Developers Keep Downloading DeBERTa-v3-small Microsoft Research’s DeBERTa-v3-small drew about 777,000 Hugging Face downloads in the reported month.
Takeaways − DeBERTa-v3-small is a 44M-parameter English encoder with a 128K vocabulary from Microsoft Research. Uses ELECTRA-style replaced-token-detection training instead of masked language modeling for better sample efficiency. Introduces Gradient-Disentangled Embedding Sharing to kill the generator/discriminator tug-of-war on shared embeddings. Scores 88.3 on MNLI-m and 82.8 F1 on SQuAD 2.0, matching models roughly twice its size. Popular backbone for classification, NLI, prompt-injection detection, safety filters, and lightweight rerankers. English-only and encoder-only; use mDeBERTa or decoder LLMs for generation or multilingual tasks. Why Developers Keep Downloading DeBERTa-v3-small Microsoft Research’s DeBERTa-v3-small drew about 777,000 Hugging Face downloads in the reported month. At the time of writing, the Hub also listed 213 fine-tunes and 20 quantized derivatives. Those changing figures measure artifact pulls, including automated jobs, but they show sustained activity around a model released before the current wave of decoder-only LLMs. The appeal comes from a practical combination: competitive natural-language understanding scores, six transformer layers, broad framework support, and an MIT license. Its unusually large vocabulary complicates the size story, however. The transformer backbone has 44 million parameters, while the complete base checkpoint contains roughly 142 million. The model card’s 44 million figure covers the transformer backbone and excludes the token embedding table. A 128,000-token vocabulary with 768-dimensional embeddings adds about 98 million parameters, accounting for roughly 69% of the checkpoint. Microsoft’s model card reports the following results. The parameter column covers each transformer backbone, keeping the comparison separate from vocabulary-dependent embedding sizes. You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.
Sources
Related stories

One year in: How Microsoft Research Asia – Singapore is advancing research, partnership and talent for real-world impact
On July 24, 2025, Microsoft Research Asia – Singapore (MSRA – Singapore) opened its doors as Microsoft’s first research lab in Southeast Asia.

Broadening access to Skala creates a faster path to predictive DFT
At a glance Skala 1.1 demonstrates the continuously improving nature of Microsoft Research’s deep-learning DFT approach: trained on 2.5× more data than…

Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement
Research Note: CARE-X is a research model and not a Microsoft product offering or medical device.

Teaching AI to speak the language of pathology
Artificial intelligence is already helping clinicians spot patterns in diagnostic images and detect signs of disease. But many of today’s pathology AI systems are built for a single task and must be rebuilt for each new application.