Can you fine-tune on localized steering of an LLM?

hok@lemmy.dbzer0.com · edit-2 12 days ago

Can you fine-tune on localized steering of an LLM?

hok@lemmy.dbzer0.com · edit-2 6 days ago

Thanks for your answer. I think to be clear, what I’m looking for is a kind of masked fine-tuning. You see, I want to “steer” a particular output instead of providing complete examples, which are costly to create.

The steering would be something like this:

I have an LLM generate a sequence.
I find exactly where the LLM goes “off track” and correct it there (for only maybe 10-20 tokens instead of correcting the rest of the generation manually).
The LLM continues “on track” until it goes off track again.

What I would like to do is train the model based on these corrections I give it, where many corrections might be part of the same overall generation. Conceptually I think each correction must have some training value. I don’t know much about masking, but what I mean here is that I don’t want it to train on a few tens or hundreds of (incomplete) samples but rather thousands of (masked) “steers” that correct the course of the rest of the sample’s generated text.