• Live Crypto Prices
  • Crypto News
    • Worldwide
      • Bitcoin
      • Ethereum
      • Altcoin
      • Blockchain
      • Regulation
    • Australian Crypto News
  • Education
    • Cryptocurrency For Beginners
    • Where to Buy Cryptocurrency
    • Where to Store Cryptos
    • Cryptocurrency Tax in Australia 2021
No Result
View All Result
CryptoABC.net
No Result
View All Result

NVIDIA Unveils Mistral-NeMo-Minitron 8B Model with Superior Accuracy

August 22, 2024
in Blockchain
Reading Time: 3min read
0 0
A A
0
Nvidia Plans to add Innovation in the Metaverse with Software, Marketplace Deals
0
SHARES
41
VIEWS
ShareShareShareShareShare


Tony Kim
Aug 22, 2024 05:37

NVIDIA’s new Mistral-NeMo-Minitron 8B model demonstrates superior accuracy across nine benchmarks, utilizing advanced pruning and distillation techniques.





NVIDIA, in collaboration with Mistral AI, has announced the release of the Mistral-NeMo-Minitron 8B model, a highly advanced open-access large language model (LLM). According to the NVIDIA Technical Blog, this model surpasses other models of a similar size in terms of accuracy on nine popular benchmarks.

Advanced Model Pruning and Distillation

The Mistral-NeMo-Minitron 8B model was developed by width-pruning the larger Mistral NeMo 12B model, followed by a light retraining process using knowledge distillation. This methodology, originally proposed by NVIDIA in their paper on Compact Language Models via Pruning and Knowledge Distillation, has been validated through multiple successful implementations, including the NVIDIA Minitron 8B and 4B models, as well as the Llama-3.1-Minitron 4B model.

Model pruning involves reducing the size and complexity of a model by either dropping layers (depth pruning) or neurons and attention heads (width pruning). This process is often paired with retraining to recover any lost accuracy. Model distillation, on the other hand, transfers knowledge from a large, complex model (the teacher model) to a smaller, simpler model (the student model), aiming to retain much of the predictive power of the original model while being more efficient.

The combination of pruning and distillation allows for the creation of progressively smaller models from a large pretrained model. This approach significantly reduces the computational cost, as only 100-400 billion tokens are needed for retraining, compared to the much larger datasets required for training from scratch.

Mistral-NeMo-Minitron 8B Performance

The Mistral-NeMo-Minitron 8B model demonstrates leading accuracy on several benchmarks, outperforming other models in its class, including the Llama 3.1 8B and Gemma 7B models. The table below highlights the performance metrics:








 Training tokensWino-Grande 5-shotARC Challenge 25-shotMMLU 5-shotHella Swag 10-shotGSM8K 5-shotTruthfulQA 0-shotXLSum en (20%) 3-shotMBPP 0-shotHuman Eval 0-shot
Llama 3.1 8B15T77.2757.9465.2881.8048.6045.0630.0542.2724.76
Gemma 7B6T786164825045173932
Mistral-NeMo-Minitron 8B380B80.3564.4269.5183.0358.4547.5631.9443.7736.22
Mistral NeMo 12BN/A82.2465.1068.9985.1656.4149.7933.4342.6323.78

Table 1. Accuracy of the Mistral-NeMo-Minitron 8B base model compared to the teacher Mistral-NeMo 12B, Gemma 7B, and Llama-3.1 8B base models. Bold numbers represent the best among the 8B model class

Implementation and Future Work

Following the best practices of structured weight pruning and knowledge distillation, the Mistral-NeMo 12B model was width-pruned to yield the 8B target model. The process involved fine-tuning the unpruned Mistral NeMo 12B model using 127 billion tokens to correct for distribution shifts, followed by width-only pruning and distillation using 380 billion tokens.

The Mistral-NeMo-Minitron 8B model showcases superior performance and efficiency, making it a significant advancement in the field of AI. NVIDIA plans to continue refining the distillation process to produce even smaller and more accurate models. The implementation of this technique will be gradually integrated into the NVIDIA NeMo framework for generative AI.

For further details, visit the NVIDIA Technical Blog.

Image source: Shutterstock


Credit: Source link

ShareTweetSendPinShare
Previous Post

Historical Data Suggests Bitcoin Could Rise 1,000%, Here’s Why

Next Post

Ethereum Is Flat, and Whales Selling: More Pain to Follow?

Next Post
Ethereum Is Flat, and Whales Selling: More Pain to Follow?

Ethereum Is Flat, and Whales Selling: More Pain to Follow?

You might also like

Google Gemini AI Predicts Jaw-Dropping Micron Technology Stock Price by End of 2026

Google Gemini AI Predicts Jaw-Dropping Micron Technology Stock Price by End of 2026

June 25, 2026
Vitalik Buterin Unveils 40% Ethereum Foundation Budget Cut in Push for Leaner Future

Vitalik Buterin Unveils 40% Ethereum Foundation Budget Cut in Push for Leaner Future

June 24, 2026
XRP Prepares for July Bounce-Back as Price History Points to

XRP Prepares for July Bounce-Back as Price History Points to

June 27, 2026
BitGo Implements 15% Workforce Reduction In Shift To AI Infrastructure

BitGo Implements 15% Workforce Reduction In Shift To AI Infrastructure

June 26, 2026
Kalshi Shows 69% Odds Bitcoin Hits $50,000 Before $100,000

Bitcoin 25-Delta Put-Call Skew Widens Amid Market Consolidation

June 26, 2026
BOJ deputy warns on inflation as Polymarket puts 2026 Fed hike odds at 66%

May inflation hits 4.1% as Polymarket sees 79% odds of zero Fed cuts in 2026

June 26, 2026
CryptoABC.net

This is an Australian online news/education portal that aims to provide the latest crypto news, real-time updates, education and reviews within Australia and around the world. Feel free to get in touch with us!

What's New Here!

Bitcoin Addresses Holding Between 100 and 10,000 BTC Hit a 7-Week High

Tokenized Deposits Gain Traction as Banks Race to Build

June 29, 2026
Bitcoin Defends $59K Support as Q2 Closes With Rare Back-to-

CryptoQuant Flags Rising Bitcoin Whale Share On Gate As BTC Holds Below $60,000

June 29, 2026

Subscribe Now

  • Contact Us
  • Privacy Policy
  • Terms of Use
  • DMCA

© 2021 cryptoabc.net - All rights reserved!

No Result
View All Result
  • Live Crypto Prices
  • Crypto News
    • Worldwide
      • Bitcoin
      • Ethereum
      • Altcoin
      • Blockchain
      • Regulation
    • Australian Crypto News
  • Education
    • Cryptocurrency For Beginners
    • Where to Buy Cryptocurrency
    • Where to Store Cryptos
    • Cryptocurrency Tax in Australia 2021

© 2021 cryptoabc.net - All rights reserved!

Welcome Back!

Login to your account below

Forgotten Password?

Create New Account!

Fill the forms below to register

All fields are required. Log In

Retrieve your password

Please enter your username or email address to reset your password.

Log In
Please enter CoinGecko Free Api Key to get this plugin works.