Skip to content

EthenEthenEthen

Open Source Model Profile · MegaScience

Qwen2.5-3B-MegaScience

Qwen2.5-3B-MegaScience is a 3.09B-parameter Qwen2 text-generation model from MegaScience. The model card describes it as one of the models trained in the MegaScience project on post-training datasets for science reasoning.

Publisher
MegaScience
Task
text-generation
Model type
qwen2
License
apache-2.0
Library
transformers
Publication status
Approved for indexing

Model overview

MegaScience publishes Qwen2.5-3B-MegaScience as a text-generation model. Captured config identifies Qwen2ForCausalLM with a qwen2 model type, and Safetensors metadata reports 3,085,938,688 parameters. Card data records apache-2.0. Hub tags list Qwen/Qwen2.5-3B as base model and finetune, and the captured model card's model-tree block names the same base.

Recorded capabilities

About 3.09B parameters

Safetensors metadata reports 3,085,938,688 parameters, or about 3.09B.

Apache-2.0 licensing

Captured metadata records an apache-2.0 license.

MegaScience project card

The model card describes this repository as a MegaScience project model trained on post-training datasets for science reasoning, and names MegaScience/MegaScience as the dataset used to train it.

Documented Transformers chat usage

The model card shows AutoTokenizer and AutoModelForCausalLM loading with apply_chat_template.

Use cases in the source record

  • Transformers text-generation and chat-template workflows following the card's AutoModelForCausalLM example.
  • Experiments described by the publisher as part of MegaScience post-training work on science-reasoning datasets.

Limitations and unknowns

  • No evaluation scores were extracted; the card includes an Evaluation Results heading without captured metrics.
  • Provider state is historical snapshot data, not independently refreshed current availability.
  • Training recipe values, dataset assignment, and the Qwen/Qwen2.5-3B base-model listing come from hub tags and the publisher model card, not independently verified Ethen measurements. Max length 4,096 is a training-recipe field, not an independently verified context window.

Source and provenance

Source: MegaScience/Qwen2.5-3B-MegaScience

Captured: Unknown. Processed: 2026-09-07T19:35:41.739349+00:00.

MegaScience: Pushing the Frontiers of Post-Training Datasets for Science Reasoning This repository contains the Qwen2.5-3B-MegaScience model, one of the models trained as part of the MegaScience project. For the official code, data processing pipeline, and evaluation system, please refer to the MegaScience GitHub repository . Qwen2.5-3B-MegaScience Usage You can use this model with the Hugging Face transformers library: from transformers import AutoModelForCausalLM, AutoTokenizer model_name = "MegaScience/Qwen2.5-3B-MegaScience" tokenizer = AutoTokenizer.from_pretrained(model_name) model = AutoModelForCausalLM.from_pretrained(model_…

F001F002F003F004F005F006F007F009F010F011F014F015F016F017F018