wandb
/

gemma-2b-zephyr-dpo

Text Generation

text-generation-inference

Inference Endpoints

Model card Files Files and versions Community

Gemma 2B Zephyr DPO

The Zephyr DPO recipe applied on top of SFT finetuned Gemma 2B

Model description

Model type: A 8.5B parameter GPT-like model fine-tuned on a mix of publicly available, synthetic datasets.
Language(s) (NLP): Primarily English
Finetuned from model: wandb/gemma-2b-zephyr-sft

Recipe

We trained using the DPO script in alignment handbook recipe and logging to W&B

Visit the W&B workspace here

License

This model has the same license as the original Gemma model collection

Compute provided by Lambda Labs - 8xA100 80GB node

around 13 hours of training

Downloads last month: 12

Safetensors

Model size

2.51B params

Tensor type

BF16

·

Inference Providers NEW

Text Generation

This model is not currently available via any of the supported third-party Inference Providers, and the model is not deployed on the HF Inference API.

Model tree for wandb/gemma-2b-zephyr-dpo

Base model

google/gemma-2b

Finetuned

wandb/gemma-2b-zephyr-sft

Finetuned

(7)

this model

Merges

1 model

Dataset used to train wandb/gemma-2b-zephyr-dpo

Spaces using wandb/gemma-2b-zephyr-dpo 2