ToolAlpaca unlearning: ToolDelete-DPO (paper checkpoint)

ToolDelete-DPO model continued to epoch 7. This checkpoint matches the ToolAlpaca paper-table results below, evaluated on 100 forget and 100 retain tasks with GPT-3.5-turbo as simulator and GPT-4-0613 as judge. The assisted condition uses the original ToolAlpaca-13B Thought planner.

Condition Forget process correctness Retain process correctness
Default 45% 55%
Assisted 53% 64%

Paired harness recovery rate: 10/55 (18.18%).

Weight SHA-256: 59dbd76f187fe619c0e4c55640d8db0df508c559a67ad1d677d7943c44b5aa99.

Downloads last month
236
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for OptimAI-Lab/ToolAlpaca_unlearn_ToolDelete-DPO

Finetuned
(6)
this model

Collection including OptimAI-Lab/ToolAlpaca_unlearn_ToolDelete-DPO