AI & Tools / Field note
Choose an AI model by testing the workflow it needs to support
Build a practical evaluation around your own inputs, review criteria and operational constraints before comparing model configurations.
The most useful model comparison begins with the work, rather than a general ranking. A system that checks product records needs different evidence of quality from one that drafts support replies or reviews images. A compelling demonstration on an unrelated task does not settle that choice.
Treat the model as one part of a complete configuration: the instructions, input preparation, available tools, output format and review process. That is the system your users will experience, and it is the system worth testing.
Define success before comparing candidates
Write down what a correct result must contain and what would make it unacceptable. For a catalogue review, that might mean identifying the right record, citing the relevant field and leaving an unsupported correction unresolved.
For a support draft, the criteria might include answering the actual question, using the supplied policy accurately and avoiding promises that the available record cannot support. These are different tests, even if both workflows produce text.
OpenAI's evaluation guidance recommends defining an objective, collecting a relevant dataset, establishing metrics and comparing results as the system changes. It also warns against evaluation data that does not reflect the intended use. OpenAI evaluation best practices.
Use examples that reveal the difficult cases
Build a small, reviewed set of representative inputs. Include straightforward examples, incomplete information, contradictory records and cases where the correct outcome is to ask for more context.
Keep some examples aside while you refine the instructions. Otherwise, it is easy to tune the prompt around the same familiar cases and mistake that improvement for broader reliability.
Record the expected behaviour in plain language. You do not always need one exact reference answer, especially for writing tasks. You do need enough clarity for two reviewers to recognise the difference between a useful response and an unsupported one.
Compare actual configurations
Keep the input data and success criteria consistent, then record the model identifier, settings and prompt version used for each run. Product names alone are too broad to describe a reproducible test.
If one candidate receives different tools, a longer source document or a different review process, record that difference. The comparison may still be useful, but the result describes the whole configuration rather than an isolated model advantage.
Look at recurring failure types alongside overall quality. A system that usually writes polished responses but occasionally substitutes the wrong product identifier may be unsuitable for the task without an additional check.
Include the cost of the complete task
Measure the time and resources needed to reach an acceptable result, including retries and human editing. A lower request cost does not automatically make the finished workflow less expensive to operate, and a slower response may be acceptable in a background review queue.
Check current provider pricing and feature availability when implementing the chosen configuration. Avoid building a long-lived decision around a price or capability remembered from an earlier release.
Choose a configuration that meets the defined quality threshold and fits the operational constraints. Keep the evaluation set so future model, prompt or integration changes can be checked against the same task. The decision becomes easier to explain when it rests on examples from the work your system actually needs to do.