Wednesday, September 9, 2026
Social icon element need JNews Essential plugin to be activated.
No Result
View All Result
Digital Currency Pulse
  • Home
  • Crypto/Coins
  • NFT
  • AI
  • Blockchain
  • Metaverse
  • Web3
  • Exchanges
  • DeFi
  • Scam Alert
  • Analysis
Crypto Marketcap
Digital Currency Pulse
  • Home
  • Crypto/Coins
  • NFT
  • AI
  • Blockchain
  • Metaverse
  • Web3
  • Exchanges
  • DeFi
  • Scam Alert
  • Analysis
No Result
View All Result
Digital Currency Pulse
No Result
View All Result

Curiosity-Driven Reinforcement Learning from Human Feedback CD-RLHF: An AI Framework that Mitigates the Diversity Alignment Trade-off In Language Models

January 31, 2025
in Artificial Intelligence
Reading Time: 4 mins read
A A
0

[ad_1]

Massive Language Fashions (LLMs) have change into more and more reliant on Reinforcement Studying from Human Suggestions (RLHF) for fine-tuning throughout varied purposes, together with code era, mathematical reasoning, and dialogue help. Nevertheless, a big problem has emerged within the type of lowered output range when utilizing RLHF. Analysis has recognized a essential trade-off between alignment high quality and output range in RLHF-trained fashions. When these fashions align extremely with desired goals, they present restricted output variability. This limitation poses issues for inventive open-ended duties equivalent to story era, knowledge synthesis, and red-teaming, the place numerous outputs are important for efficient efficiency.

Present approaches to LLM alignment have centered on enhancing instruction following, security, and reliability by means of RLHF, however these enhancements typically come at the price of output range. Varied strategies have been developed to handle this problem, together with the usage of f-divergence with DPO/PPO algorithms, which try and stability range and alignment. Different approaches combine analysis metrics like SelfBLEU and Sentence-BERT into RL fine-tuning to spice up range, significantly for red-teaming duties. Furthermore, some researchers have explored curiosity-driven reinforcement studying strategies, starting from count-based approaches to prediction error-based strategies. Regardless of these efforts, the elemental trade-off between alignment high quality and output range stays a big problem.

Researchers from Baidu have proposed a novel framework referred to as Curiosity-driven Reinforcement Studying from Human Suggestions (CD-RLHF) to handle the diversity-alignment trade-off in language fashions. This method incorporates curiosity as an intrinsic reward mechanism throughout the RLHF coaching stage, working alongside conventional extrinsic rewards from the reward mannequin. CD-RLHF makes use of ahead dynamics to compute prediction errors of state representations, which helps estimate curiosity ranges. A key function of this method is that steadily visited states step by step change into much less fascinating to the mannequin. This twin reward system goals to keep up excessive alignment high quality whereas selling numerous outputs by means of different token decisions at every choice level.

The implementation and analysis of CD-RLHF encompasses a number of parts and datasets. The structure was examined on two main datasets: TL;DR for textual content summarization, containing 93k human-annotated desire pairs, and UltraFeedback for instruction following, with 61.1k coaching pairs. The framework was applied utilizing varied base fashions together with Gemma-2B, Gemma-7B, Llama-3.2-1B, and Llama-3.2-3B, all educated throughout the DeepSpeed-Chat framework. The coaching knowledge was distributed throughout SFT, RM, and PPO levels in a 20/40/40 ratio. For comparability, baseline strategies together with vanilla RLHF and Despatched-Rewards are applied, which use SelfBLEU and Sentence-BERT scores as further rewards throughout coaching.

The experimental outcomes reveal CD-RLHF’s superior efficiency throughout a number of analysis metrics and fashions. Within the TL;DR summarization process, CD-RLHF achieves important enhancements in output range displaying positive factors of 16.66% and 6.22% on Gemma-2B and Gemma-7B respectively in comparison with the RLHF baseline. For the UltraFeedback instruction-following process, the tactic reveals much more spectacular outcomes, with range enhancements starting from 7.35% to 14.29% throughout totally different fashions whereas sustaining robust alignment high quality. Exterior validation by means of GPT-4 analysis confirmed CD-RLHF attaining win charges of as much as 58% towards the PPO baseline on TL;DR and a median of 62% on UltraFeedback.

In conclusion, researchers launched CD-RLHF which represents a big development in addressing the diversity-alignment trade-off in language mannequin coaching. The framework combines curiosity-driven exploration with conventional extrinsic rewards to reinforce output range whereas sustaining alignment high quality, as proven by means of in depth testing on TL;DR summarization and UltraFeedback instruction-following duties. Regardless of these achievements, a number of challenges stay, together with the necessity to stability totally different reward scales and the persistent hole between the output range of SFT, and RLHF-trained fashions. Whereas CD-RLHF mitigates the trade-off between range and alignment, additional analysis is required to totally bridge this hole and obtain optimum efficiency throughout each metrics.

Try the Paper and GitHub Web page. All credit score for this analysis goes to the researchers of this undertaking. Additionally, don’t overlook to comply with us on Twitter and be part of our Telegram Channel and LinkedIn Group. Don’t Overlook to hitch our 70k+ ML SubReddit.

🚨 Meet IntellAgent: An Open-Supply Multi-Agent Framework to Consider Advanced Conversational AI System (Promoted)

Sajjad Ansari is a closing 12 months undergraduate from IIT Kharagpur. As a Tech fanatic, he delves into the sensible purposes of AI with a deal with understanding the impression of AI applied sciences and their real-world implications. He goals to articulate advanced AI ideas in a transparent and accessible method.

✅ [Recommended] Be a part of Our Telegram Channel

[ad_2]

Source link

Tags: AlignmentCDRLHFCuriosityDrivenDiversityFeedbackFrameworkhumanlanguageLearningMitigatesmodelsReinforcementTradeoff
Previous Post

Everything You Need to Know About Real-Time 3D Experiences

Next Post

Weekly Roundup: “Gritty” FriendsCaps Available Now, Comic Book #1 Inside Look, New Mint Mink Partnership… and MORE! | by VeeFriends | Jan, 2025

Next Post
Weekly Roundup: “Gritty” FriendsCaps Available Now, Comic Book #1 Inside Look, New Mint Mink Partnership… and MORE! | by VeeFriends | Jan, 2025

Weekly Roundup: “Gritty” FriendsCaps Available Now, Comic Book #1 Inside Look, New Mint Mink Partnership… and MORE! | by VeeFriends | Jan, 2025

Trump Effect? Solana Stablecoin Supply Jumps 73% Since Mid-January

Trump Effect? Solana Stablecoin Supply Jumps 73% Since Mid-January

Grayscale Debuts New Dogecoin Trust—And Files to Convert It Into ETF

Grayscale Debuts New Dogecoin Trust—And Files to Convert It Into ETF

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Social icon element need JNews Essential plugin to be activated.

CATEGORIES

  • Analysis
  • Artificial Intelligence
  • Blockchain
  • Crypto/Coins
  • DeFi
  • Exchanges
  • Metaverse
  • NFT
  • Scam Alert
  • Web3
No Result
View All Result

SITEMAP

  • About us
  • Disclaimer
  • DMCA
  • Privacy Policy
  • Terms and Conditions
  • Cookie Privacy Policy
  • Contact us

Copyright © 2024 Digital Currency Pulse.
Digital Currency Pulse is not responsible for the content of external sites.

No Result
View All Result
  • Home
  • Crypto/Coins
  • NFT
  • AI
  • Blockchain
  • Metaverse
  • Web3
  • Exchanges
  • DeFi
  • Scam Alert
  • Analysis
Crypto Marketcap

Copyright © 2024 Digital Currency Pulse.
Digital Currency Pulse is not responsible for the content of external sites.