Skywork has launched the second-generation reward models, Skywork-Reward-V2, continuing its leadership in open-source AI innovation. Building on the success of the original Skywork-Reward series released in September 2024, which garnered over 750,000 downloads on HuggingFace and powered top results on benchmarks like RewardBench, the V2 series introduces eight new models ranging from 600M to 8B parameters, achieving state-of-the-art rankings across seven major evaluation benchmarks.
These models are available for download here:
Skywork-SynPref-40M: Hybrid Human-Machine Data for Precision Reward Modeling
Central to the success of Skywork-Reward-V2 is Skywork-SynPref-40M, a groundbreaking hybrid dataset of 40 million preference pairs. Created through a two-stage human-machine collaboration process, it blends rigorously verified human annotations with large-scale automated expansions via Large Language Models. The process ensures both high quality and broad coverage, with 26 million samples surviving rigorous screening for optimal performance.
Stage 1 involved building a gold standard dataset through strict human verification and leveraging LLMs to generate high-quality silver standard data, followed by iterative training cycles to identify weaknesses and refine data. Stage 2 scaled the process through fully automated filtering and re-annotation, significantly reducing manual workload while maintaining precision.
Matching Large-Scale Model Performance with Smaller Models
The V2 series demonstrates that smaller models can match or surpass the performance of much larger counterparts when trained on rich, high-quality datasets. For example, the 600M-parameter Skywork-Reward-V2-Qwen3-0.6B nearly matches the results of the previous generation’s 27B-parameter flagship, while the largest model, Skywork-Reward-V2-Llama-3.1-8B, leads all mainstream benchmarks.
Breakthroughs in Multi-Dimensional Human Preference Alignment
Skywork-Reward-V2 excels across multiple capability dimensions, including Best-of-N scaling, bias resistance, complex instruction handling, and truthfulness assessment. This broad generalization makes it applicable to a wide range of RLHF and RLVR tasks, from natural language reasoning to mathematics and programming.
Impact and Future Vision
This release represents a significant step toward building the foundation for future AI infrastructure. Reward models are increasingly seen not just as evaluators, but as the “compass” for intelligent systems—guiding alignment with human values and enabling continuous evolution toward meaningful objectives.
Skywork plans to expand its research into alternative training techniques, modeling objectives, and unified reward systems, with the vision that such systems will become central to large-scale AI training pipelines.
In addition to the V2 series, Skywork introduced the world’s first deep research AI workspace agents in May 2025, now available for exploration at skywork.ai.
For the latest in AI, IoT, cybersecurity, and tech research insights, visit ITech360hub.
