Alpaca-tuned Llama-3 8B
According to the model card, this fine-tunes AmberYifan/llama3-8b-full-pretrain-junk-tweet-1m-en on the alpaca_en dataset.
Open Source Model Profile · AmberYifan
llama3-8b-full-pretrain-junk-tweet-1m-en-sft is an 8.03B-parameter Llama fine-tune from AmberYifan. According to the model card, it tunes the junk-tweet base on alpaca_en under a llama3 license.
llama3-8b-full-pretrain-junk-tweet-1m-en-sft is published by AmberYifan as a Transformers text-generation fine-tune. The captured configuration identifies LlamaForCausalLM with model type llama and about 8.03B Safetensors parameters, and card data records llama3. According to the model card, it fine-tunes AmberYifan/llama3-8b-full-pretrain-junk-tweet-1m-en on alpaca_en with a documented 3-epoch schedule.
According to the model card, this fine-tunes AmberYifan/llama3-8b-full-pretrain-junk-tweet-1m-en on the alpaca_en dataset.
According to the model card, training ran for 3.0 epochs with cosine scheduling, learning rate 1e-05, and multi-GPU batching.
Captured configuration identifies LlamaForCausalLM with 8,030,261,248 Safetensors parameters, or about 8.03B.
Card data records a llama3 license for this fine-tune.
Source: AmberYifan/llama3-8b-full-pretrain-junk-tweet-1m-en-sft
Captured: Unknown. Processed: 2026-09-07T19:35:33.816834+00:00.
llama3-8b-full-pretrain-junk-tweet-1m-en-sft This model is a fine-tuned version of AmberYifan/llama3-8b-full-pretrain-junk-tweet-1m-en on the alpaca_en dataset. Model description More information needed Intended uses & limitations More information needed Training and evaluation data More information needed Training procedure Training hyperparameters The following hyperparameters were used during training: learning_rate: 1e-05 train_batch_size: 1 eval_batch_size: 8 seed: 42 distributed_type: multi-GPU num_devices: 8 gradient_accumulation_steps: 2 total_train_batch_size: 16 total_eval_batch_size: 64 optimizer: Use adamw_torch with bet…
F001F002F003F004F005F006F007F009F010F011F012