Skip to content

EthenEthenEthen

Open Source Model Profile · AmberYifan

llama3-8b-full-pretrain-junk-tweet-1m-en-sft

llama3-8b-full-pretrain-junk-tweet-1m-en-sft is an 8.03B-parameter Llama fine-tune from AmberYifan. According to the model card, it tunes the junk-tweet base on alpaca_en under a llama3 license.

Publisher
AmberYifan
Task
text-generation
Model type
llama
License
llama3
Library
transformers
Publication status
Accepted · not indexed

Model overview

llama3-8b-full-pretrain-junk-tweet-1m-en-sft is published by AmberYifan as a Transformers text-generation fine-tune. The captured configuration identifies LlamaForCausalLM with model type llama and about 8.03B Safetensors parameters, and card data records llama3. According to the model card, it fine-tunes AmberYifan/llama3-8b-full-pretrain-junk-tweet-1m-en on alpaca_en with a documented 3-epoch schedule.

Recorded capabilities

Alpaca-tuned Llama-3 8B

According to the model card, this fine-tunes AmberYifan/llama3-8b-full-pretrain-junk-tweet-1m-en on the alpaca_en dataset.

Documented 3-epoch schedule

According to the model card, training ran for 3.0 epochs with cosine scheduling, learning rate 1e-05, and multi-GPU batching.

8.03B Llama build

Captured configuration identifies LlamaForCausalLM with 8,030,261,248 Safetensors parameters, or about 8.03B.

Llama3 license

Card data records a llama3 license for this fine-tune.

Use cases in the source record

  • Conversational text-generation experiments using the alpaca-tuned Llama-3 8B setup.
  • Full fine-tuning reproduction work using the documented hyperparameters and llama-factory tooling.

Limitations and unknowns

  • No evaluation results were extracted from this record.
  • No context-window value was extracted from this record.
  • Intended uses, limitations, and evaluation data are marked as needing more information in the model card.
  • Provider state is historical snapshot data, not independently refreshed current availability.

Source and provenance

Source: AmberYifan/llama3-8b-full-pretrain-junk-tweet-1m-en-sft

Captured: Unknown. Processed: 2026-09-07T19:35:33.816834+00:00.

llama3-8b-full-pretrain-junk-tweet-1m-en-sft This model is a fine-tuned version of AmberYifan/llama3-8b-full-pretrain-junk-tweet-1m-en on the alpaca_en dataset. Model description More information needed Intended uses & limitations More information needed Training and evaluation data More information needed Training procedure Training hyperparameters The following hyperparameters were used during training: learning_rate: 1e-05 train_batch_size: 1 eval_batch_size: 8 seed: 42 distributed_type: multi-GPU num_devices: 8 gradient_accumulation_steps: 2 total_train_batch_size: 16 total_eval_batch_size: 64 optimizer: Use adamw_torch with bet…

F001F002F003F004F005F006F007F009F010F011F012