Paper Detail

Solar Open 2 Technical Report

Sungrae Park, Sanghoon Kim, Gyoungjin Gim, Jungho Cho, Hyunwoong Ko, Minbyul Jeong, Minjeong Kim, Keunwoo Choi, Chaehun Shin, Chanwoong Yoon, Dongjun Kim, Eunwon Kim, Gyungin Shin, Hyeonju Lee, Hyungkyu Kang, Inseo Song, Jisu Bae, Jiyoon Han, Jiyun Lee, Joonkee Kim, Junyeop Lee, Mikyoung Cha, Sangwon Yu, Sehwan Joo, Seokyoon Kang, Seonghoon Yang, Seung Shin, Seunghyun Lee, Seungseop Lim, Seungyoun Shin, Sukyung Lee, Taegyeong Eo, Taehwan Oh, Taewhoo Lee, Wonho Song, Wonjun Oh, Wonseok Hwang, Yunsu Kim, Yura Shim, Hwalsuk Lee, Sunghun Kim, Du-Seong Chang, Kyunghyun Cho, Seungju Han, Yejin Choi, Junsuk Choe, Hwaran Lee, Minjeong Ban, Yun Taewon, Hwanjun Song, Jae-Gil Lee, KyungTae Lim, Alice Oh

arxiv Score 10.6

Published 2026-07-22 · First seen 2026-07-23

General AI

Abstract

We present Solar Open 2, a 250B-A15B Mixture-of-Experts language model built for long-horizon agentic tasks, scaled up from Solar Open 1 (Solar Open 100B). To hold entire agent trajectories in a single context, Solar Open 2 reaches a 1M-token window through a hybrid attention stack that interleaves one softmax layer among every three linear-attention layers, using no positional encoding and a gated delta rule extended to negative eigenvalues. To train at this scale under a fixed compute budget, we make training efficient in two ways: a stronger starting point, and higher-value data. For the starting point, we initialize Solar Open 2 from Solar Open 1, transferring the 5.69B-parameter shared skeleton that survives the architectural change and learning everything else through full pre-training. For the data, we curate for value per token: quality- and rarity-aware data curation and mixture-ratio optimization refine a 20T pool into a 10T mixture that, at equal token budget, outperforms the Solar Open 1 recipe. To build its agent skills, we train twelve domain specialists across purpose-built scenarios, then consolidate them into a single model by Multi-teacher On-Policy Distillation (MOPD). Against comparably sized open-weight models on English benchmarks, Solar Open 2 leads on MMLU-Pro, LiveCodeBench, and the APEX-Agents agentic suite, and stays competitive with the strongest (DeepSeek-V4-Flash and MiMo-V2.5) elsewhere. On Korean benchmarks, Solar Open 2 records the highest average of any model compared, including fast-tier closed APIs, and on Ko-GDPval, an in-house Korean officework-agent benchmark, it is competitive with DeepSeek-V4-Pro (1.6T) at less than a sixth of its size.

Workflow Status

Review status
pending
Role
unreviewed
Read priority
soon
Vote
Not set.
Saved
no
Collections
Not filed yet.
Next action
Not filled yet.

Reading Brief

No structured notes yet. Add `summary_sections`, `why_relevant`, `claim_impact`, or `next_action` in `papers.jsonl` to enrich this view.

Why It Surfaced

No ranking explanation is available yet.

Tags

No tags.

BibTeX

@article{park2026solar,
  title = {Solar Open 2 Technical Report},
  author = {Sungrae Park and Sanghoon Kim and Gyoungjin Gim and Jungho Cho and Hyunwoong Ko and Minbyul Jeong and Minjeong Kim and Keunwoo Choi and Chaehun Shin and Chanwoong Yoon and Dongjun Kim and Eunwon Kim and Gyungin Shin and Hyeonju Lee and Hyungkyu Kang and Inseo Song and Jisu Bae and Jiyoon Han and Jiyun Lee and Joonkee Kim and Junyeop Lee and Mikyoung Cha and Sangwon Yu and Sehwan Joo and Seokyoon Kang and Seonghoon Yang and Seung Shin and Seunghyun Lee and Seungseop Lim and Seungyoun Shin and Sukyung Lee and Taegyeong Eo and Taehwan Oh and Taewhoo Lee and Wonho Song and Wonjun Oh and Wonseok Hwang and Yunsu Kim and Yura Shim and Hwalsuk Lee and Sunghun Kim and Du-Seong Chang and Kyunghyun Cho and Seungju Han and Yejin Choi and Junsuk Choe and Hwaran Lee and Minjeong Ban and Yun Taewon and Hwanjun Song and Jae-Gil Lee and KyungTae Lim and Alice Oh},
  year = {2026},
  abstract = {We present Solar Open 2, a 250B-A15B Mixture-of-Experts language model built for long-horizon agentic tasks, scaled up from Solar Open 1 (Solar Open 100B). To hold entire agent trajectories in a single context, Solar Open 2 reaches a 1M-token window through a hybrid attention stack that interleaves one softmax layer among every three linear-attention layers, using no positional encoding and a gated delta rule extended to negative eigenvalues. To train at this scale under a fixed compute budget, },
  url = {https://arxiv.org/abs/2607.20062},
  keywords = {cs.CL},
  eprint = {2607.20062},
  archiveprefix = {arXiv},
}

Metadata

{}