Guide to model adaptation
Instructions, knowledge or fine-tuning?
Diagnose the task before choosing the method. A missing fact, an unclear instruction and a repeated behaviour problem call for different experiments.
How do we choose between instructions, connected knowledge and fine-tuning?
Improve instructions when the task or expected response is unclear. Connect knowledge when the system needs business information at the time of a request. Consider fine-tuning when repeated behaviour or task-performance gaps remain despite suitable inputs. Compare the options on representative cases before committing to an approach.
Use the failure to choose the first test.
Assume you have a defined task, permitted inputs and reviewers who can judge the result. Use these rows to select an experiment, then check its effect.
The response misses the requested structure
First test: clarify the instructions and show suitable examples. Limit: instructions do not supply missing business records.
The response lacks current business facts
First test: retrieve the relevant approved information. Limit: irrelevant, outdated or inaccessible sources still need correction.
A recurring behaviour remains unsuitable
First test: evaluate fine-tuning with representative examples. Limit: adaptation still needs independent test cases and a useful baseline.
The process needs a business action
First test: define the tool, workflow and permission boundary. Limit: training a model does not grant authority to change records.
Keep one task constant while comparing approaches.
Illustrative task: prepare a service-response draft containing an issue summary, a reference to the relevant policy and a proposed next step. Inputs are a request and the approved service guidance. A service specialist reviews the draft before it is sent.
If the draft omits the next-step field, test clearer instructions. If it cites yesterday's policy, inspect the knowledge source and retrieval. If its task-specific behaviour remains inconsistent with the right information, compare a fine-tuned candidate with the existing approach.
Make the comparison fair and useful.
Use the same representative development cases and review criteria across candidates. Include a normal request, a changed policy, missing information and an ambiguous request. Use these cases to compare and refine approaches. Reserve a separate final test set that does not guide training, prompt changes or candidate selection, and assess the selected approach on it.
Record the model, instructions, knowledge and test versions. Check required fields, factual support and the proposed action. Compare response time, resource use and reviewer effort alongside quality. Repeat relevant cases to understand variability.
Write down the decision your evidence supports.
Complete these prompts after the comparison. A specific unanswered question is more useful than an assumed benefit.
- Observed gap: which outputs fail, and how do reviewers identify the failure?
- Chosen change: instructions, knowledge, model adaptation or workflow design?
- Evidence: which cases improved, which did not and what became worse?
- Operating effect: what changes in cost, delay, maintenance or review work?
- Decision: proceed, keep the baseline or run a named additional test?
Combine approaches when the task needs it.
A model can be adapted for a response pattern and still retrieve current policies. Keep those responsibilities distinct, so a later policy change prompts a knowledge review and a behaviour change prompts a model or instruction review.
Questions and answers
No. Retrieval supplies information with the request. Fine-tuning changes an existing model using examples. They can be used together.
Test instructions when the task definition is the problem. If required information or system access is missing, address that cause directly.
Evidence that the adapted model improves the defined task enough to justify its operating and maintenance tradeoffs. Set the acceptance criteria for your workload.
The requirement depends on the model, method and task. Assess example quality, coverage and evaluation evidence rather than adopting a universal count.