ViT pet-breed fine-tune
The model card says it fine-tunes google/vit-base-patch16-224 for image classification on the Oxford-IIIT Pet Dataset.
Open Source Model Profile · ISxOdin
vit-base-oxford-iiit-pets is an 85.8M-parameter ViT image-classification fine-tune from ISxOdin. The model card describes Oxford-IIIT pet-breed training from google/vit-base-patch16-224.
vit-base-oxford-iiit-pets is published by ISxOdin as an image-classification fine-tune. The captured configuration identifies ViTForImageClassification with a vit model type and Safetensors metadata reports 85827109 parameters. According to the model card, it fine-tunes google/vit-base-patch16-224 on the Oxford-IIIT Pet Dataset.
The model card says it fine-tunes google/vit-base-patch16-224 for image classification on the Oxford-IIIT Pet Dataset.
According to the model card, evaluation reports 0.9445 accuracy with 0.1924 loss, alongside epoch-level validation results.
According to the model card, transfer learning adapts the model to 37 cat and dog breeds with an adjusted classification head trained end-to-end.
The model card says it may not generalize beyond Oxford-IIIT breeds and that inputs should be clear, centered portraits like the training data.
Source: ISxOdin/vit-base-oxford-iiit-pets
Captured: Unknown. Processed: 2026-09-07T19:35:06.342380+00:00.
vit-base-oxford-iiit-pets This model is a fine-tuned version of google/vit-base-patch16-224 on the pcuenq/oxford-pets dataset. It achieves the following results on the evaluation set: Loss: 0.1924 Accuracy: 0.9445 Model description This model is a fine-tuned version of a pre-trained Vision Transformer ( google/vit-base-patch16-224 ) for image classification on the Oxford-IIIT Pet Dataset. It uses transfer learning to adapt a generic vision model to identify 37 different cat and dog breeds. The model head is adjusted to output the number of classes in the dataset, and it is trained end-to-end using standard classification loss. Inten…
F001F002F003F004F005F006F007F010F011F012F013F014F015F016F017F019