Change the AI.
Keep the business context.

How to compare Codex, Claude and Gemini on your actual work, without rebuilding your business knowledge every time you switch.

A DLVX field guide · 8 minute read, plus your own test · Reviewed 28 September 2026

Start with the job

The next model release is not your business strategy.

A convincing demo tells you what a tool can do under those conditions. It does not tell you whether it can find your current policy, follow your approval process, or leave the next person with a usable result.

That is why we treat tool choice as a decision we can revisit. At DLVX, shared instructions and project context sit outside individual conversations. A change of tool should not require a new explanation of the whole business.

The aim is flexibility with continuity: keep the knowledge you maintain, test the tool doing the work, and keep a record of what passed.

Your business context should survive a change of AI.

One useful distinction: Codex is a coding-agent product. Claude and Gemini name model families and associated products. Compare the actual applications, model versions, settings and permissions you would use. A product with repository access is doing a different job from a chat window with a pasted paragraph.

Different tools. The same brief.

Choose a tool, then add the business context. The example shows what needs to travel with the work. It does not simulate a vendor's performance.

Codex
Business context

Current facts · Sources and dates
Relevant decisions · Who can approve
Definition of done

“Reply to this delivery enquiry.”

The task alone does not tell the tool whether the customer is eligible, which delivery policy is current, or whether it may make a promise.

Missing: approved policy, source date, customer facts and permission to send.

Worked fictional example. No AI model runs here, and no customer information is sent anywhere.

The business brain

Keep it small enough to trust.

A central brain is a maintained set of business knowledge and working rules that different tools can access. The sharing is something you deliberately set up and test; different products do not automatically inherit each other's memory. It can begin with a small folder and a few clear records. It does not have to begin with a new software platform.

Separate enduring context from changing facts. Tone of voice may stay useful for months. Stock availability, an invoice balance or a client's latest decision needs a current source. A saved summary should help the tool find that source; it should not quietly replace it.

Keep in the shared contextMake explicit
Business identity and scopeWhat you do, who you serve, what you do not offer.
Current job and decisionsWhat changed, who decided, and where the decision came from.
Authoritative sourcesWhich record is current, who owns it and when it was checked.
Permissions and handoffWhat may be drafted, what may be changed and who reviews it.
Accepted workWhat passed, where it lives and what the next person needs.

Give a tool the relevant slice for its task, with access controls. Do not send every document to every provider by default. Credentials belong in the access system, not a shared prompt.

A lesson from our own setup

Our runtime entry instructions point different tools back to a shared operating core and scoped project context. They also distinguish stored orientation from facts that need a fresh check. The useful mechanism is the maintained handoff, not the length of the prompt.

A fair comparison

Test the job you actually want to delegate.

Pick one task with an output you can judge. Use the same input, relevant context, permission level and success criteria for each candidate. Record any tool-access differences rather than pretending they do not matter.

Begin with a small set of representative cases: a normal request, a messy one, missing information, a contradiction and something outside scope. Repeat the important cases. A handful of successes is a first screen, not a reliability estimate.

Decide what a failure means before you read the answers. An invented price, a missed source or an unauthorized action should not be averaged away by elegant writing. After those checks, compare human correction time, total turnaround, running cost and how easily the work can be continued.

Test the checker, too.

Our internal evaluation work surfaced another problem: a simplistic text check can score the wrong thing. As a fictional example, finding the word “approved” inside “not approved” is not evidence of approval. Try known good and known bad outputs before trusting a score, and check the actual result when an action matters.

OpenAI recommends task-specific evaluations and revisiting them as a system changes. Anthropic's agent guidance similarly puts clear success criteria and a simple starting design ahead of unnecessary orchestration. Our practical translation is a portable test pack you can keep even when the tool changes. OpenAI evaluation guidance; Anthropic agent guidance.

Make a comparison you can reuse.

Choose a task. Add a plain-language description without private customer data. Download a brief for the tools you want to test.

Keep names, customer records and confidential information out of this practice brief.

Your pass conditions

    Include these awkward cases

    This creates a test plan. It does not rank or run the tools.

    When the tools change

    Rerun the work, not the marketing.

    A model update, a different application, a changed tool permission or a new business policy can all change the result. Keep a dated record of the configuration that passed. Re-run the same test pack before moving consequential work, and add recent real failures to it.

    Keep the last working setup available where the product allows it. If a product does not let you pin a version, record the application and date and run lightweight checks regularly. Model availability also changes; Google's model catalogue and provider deprecation notices are a reason to maintain a portable test, not a permanent ranking. Google's model catalogue.

    Do not add a second model just to look sophisticated. It earns its place when an observed improvement is worth the extra configuration, access management and checking. One well-supported tool can be the right answer.

    Your next hour

    1. Choose a job that already happens in your business.
    2. Write down what must be true for its output to be accepted.
    3. Gather a small set of safe, representative inputs.
    4. Compare candidates and record the corrections a person had to make.

    Once you know which tool can do the job, the next question is whether the complete workflow is worth running.

    Prove the workflow
    Keep the useful part

    Take the test with you.

    Sources and what this guide claims

    Checked 28 September 2026. This is a DLVX working method, not a published head-to-head model benchmark. The interactive example is fictional. No universal ranking or routine Gemini deployment is claimed.