Qwen3-8B agent post-train
According to the model card, the model is post-trained from Qwen/Qwen3-8B with sequential SFT and RL stages for agentic tasks.
Open Source Model Profile · open-thoughts
OpenThinker-Agent-v1 is an open-thoughts 8.19B Qwen3 agent model. Its card documents Qwen3-8B post-training for terminal and coding tasks.
OpenThinker-Agent-v1 is published by open-thoughts as a text-generation model for agentic work. The captured configuration identifies Qwen3ForCausalLM with Safetensors metadata reporting 8,190,735,360 parameters. According to the model card, it is post-trained from Qwen/Qwen3-8B for tasks including Terminal-Bench 2.0 and SWE-Bench.
According to the model card, the model is post-trained from Qwen/Qwen3-8B with sequential SFT and RL stages for agentic tasks.
The model card says it is trained for agentic tasks such as Terminal-Bench 2.0 and SWE-Bench.
The card documents about 15,200 SFT traces, about 720 RL tasks, and a three-stage filtration pipeline that removes tasks with unreliable or slow verifiers.
Source: open-thoughts/OpenThinker-Agent-v1
Captured: Unknown. Processed: 2026-09-07T19:35:59.605475+00:00.
Project | SFT dataset | RL dataset | SFT model | RL model OpenThinker-Agent-v1 OpenThoughts-Agent is an open-source effort to curate the best datasets for training agents. Our first release includes datasets , models and our research codebase . OpenThinker-Agent-v1 is a model trained for agentic tasks such as Terminal-Bench 2.0 and SWE-Bench . The OpenThinker-Agent-v1 model is post-trained from Qwen/Qwen3-8B . It is SFT-ed on the OpenThoughts-Agent-v1-SFT dataset, then RL-ed on the OpenThoughts-Agent-v1-RL dataset. This model is the final model after both SFT and RL. For the model after the SFT stage only, see OpenThinker-Agent-v1-S…
F001F002F003F004F005F006F007F008F009F010F012F013F014F015F016F017