Skip to content

EthenEthenEthen

Open Source Model Profile · ISxOdin

vit-base-oxford-iiit-pets

vit-base-oxford-iiit-pets is an 85.8M-parameter ViT image-classification fine-tune from ISxOdin. The model card describes Oxford-IIIT pet-breed training from google/vit-base-patch16-224.

Publisher
ISxOdin
Task
image-classification
Model type
vit
License
apache-2.0
Library
transformers
Publication status
Accepted · not indexed

Model overview

vit-base-oxford-iiit-pets is published by ISxOdin as an image-classification fine-tune. The captured configuration identifies ViTForImageClassification with a vit model type and Safetensors metadata reports 85827109 parameters. According to the model card, it fine-tunes google/vit-base-patch16-224 on the Oxford-IIIT Pet Dataset.

Recorded capabilities

ViT pet-breed fine-tune

The model card says it fine-tunes google/vit-base-patch16-224 for image classification on the Oxford-IIIT Pet Dataset.

Publisher-reported accuracy

According to the model card, evaluation reports 0.9445 accuracy with 0.1924 loss, alongside epoch-level validation results.

37-breed classification head

According to the model card, transfer learning adapts the model to 37 cat and dog breeds with an adjusted classification head trained end-to-end.

Documented limitations

The model card says it may not generalize beyond Oxford-IIIT breeds and that inputs should be clear, centered portraits like the training data.

Use cases in the source record

  • Pet-breed image classification across the 37 documented Oxford-IIIT cat and dog breeds.
  • Portrait-style pet photo classification where inputs are clear and centered like the training data.

Limitations and unknowns

  • According to the model card, the model may not generalize well to breeds outside the Oxford-IIIT dataset.
  • According to the model card, inputs should be clear, centered, and close in style to cropped pet portraits.
  • Provider state is historical snapshot data, not independently refreshed current availability.
  • No independent Ethen evaluation results were extracted; accuracy figures come from the publisher model card.

Source and provenance

Source: ISxOdin/vit-base-oxford-iiit-pets

Captured: Unknown. Processed: 2026-09-07T19:35:06.342380+00:00.

vit-base-oxford-iiit-pets This model is a fine-tuned version of google/vit-base-patch16-224 on the pcuenq/oxford-pets dataset. It achieves the following results on the evaluation set: Loss: 0.1924 Accuracy: 0.9445 Model description This model is a fine-tuned version of a pre-trained Vision Transformer ( google/vit-base-patch16-224 ) for image classification on the Oxford-IIIT Pet Dataset. It uses transfer learning to adapt a generic vision model to identify 37 different cat and dog breeds. The model head is adjusted to output the number of classes in the dataset, and it is trained end-to-end using standard classification loss. Inten…

F001F002F003F004F005F006F007F010F011F012F013F014F015F016F017F019