Choose Your First Useful AI Task
“Use AI” is not a useful task description. A small business needs to identify a repetitive piece of work with a clear input, an observable output and a person who can check that output. Start where a mistake can be caught before it reaches a customer or changes a record. A modest trial with non-sensitive examples can show whether the tool actually reduces total effort. The first success may be a better draft or a quicker summary, not an autonomous system making decisions on the business's behalf.
Compare three candidate tasks
Drafting: a consultant might use a tool to propose a first version of a routine meeting summary or a standard response to a general enquiry. The person who attended the meeting checks facts, tone, promises and next steps before sending. Drafting works best when the reviewer knows what good output looks like. Summarising: the same person might turn their own non-sensitive project notes into a short action list, then compare every action against the originals. Classification: a shop might group generic incoming requests by topic for manual routing, without letting a prediction decide a refund, a contract or a customer's eligibility.
Each type has different failure modes. Drafts can invent details; summaries can omit an exception; classifications can quietly put an unusual case in the wrong category. List the consequence of each error and the effort required to spot it. Pick a task with enough repetition that a small time saving could matter, but limited enough that every output can be reviewed during the test. NIST's AI Risk Management Framework calls for defining the context of use, measurement and human oversight. That is a practical starting point even for a business with no technical team.
Check the data before trying a tool
Look at the input a typical case would require. Does it include personal details, confidential customer plans, employee information, payment data, unpublished commercial terms or protected intellectual property? A public tool may handle entered information under terms that do not meet the business's obligations or client promises. Do not paste live customer material merely to see what happens. Start with invented or properly anonymised test cases and inspect the provider's current terms, data retention, training settings, access controls and deletion options. Local privacy and professional rules vary; seek qualified advice when you are unsure.
The UK Information Commissioner's Office notes that generative AI used with personal data raises data protection questions, while the UK National Cyber Security Centre advises integrating security into AI workflows from the start. Neither source can determine whether a specific product is suitable for your business. You must check the exact service, the data, your contract and the jurisdiction in which you operate. If the test requires sensitive material to be useful, pause until the permissions and protections are understood.
Design a small, reviewable trial
Write one sentence describing the input, proposed output and responsible reviewer. For example: “From a short, non-confidential meeting note, draft a list of actions for the consultant to check before it is shared.” Collect ten representative cases, including a few with ambiguous phrasing. Time the current manual process on five cases. Run the tool on the other five or, preferably, apply both methods to comparable cases and include prompting, reviewing, correcting and copying in total time. Record the tool's charge per case and the hours needed to set up and administer access. Do not compare a polished manual result with an unchecked draft.
Create a short review checklist: Does each name or date appear in the source? Is any action missing? Has the tool added a promise nobody made? Is the language appropriate for the recipient? Can the reviewer trace the answer back to the input? Decide that nothing is sent or entered into a customer record until a person completes the check. Keep the test outputs separate from live operations and delete them under the policy you chose. If results are poor, changing the prompt may help, but repeated repair is also evidence that the task or tool is a poor fit.
A consultant-office example
Assume a two-person advisory firm spends twelve minutes turning a typical internal project note into an action list. It tests an AI-assisted draft on ten synthetic or approved non-sensitive notes. Prompts take two minutes, review and corrections seven minutes, and copying to the agreed record one minute, for ten minutes total per case. That is a two-minute saving before subscription fees and setup time. If the tool costs 30 currency units per month and the firm handles 20 similar notes, the apparent 40 minutes saved must be compared with the 30-unit fee and the value of staff time. These are illustrative assumptions, not a forecast of any product's performance.
Suppose three of the ten outputs omit a deadline that the reviewer must recover from the note. Record a 30% correction incidence for that small sample, with the exact error type. If a deadline can be missed in live work, the review checklist must reliably catch it; otherwise the task is unsuitable for this use. A second trial might use a stricter format and then measure again. If total review time rises above the manual twelve minutes, stop the pilot even if the first draft looks impressive. Count quality as well as speed.
Decide whether to keep, change or drop it
Track three numbers: total minutes per case including review, the proportion of outputs requiring a material correction, and total tool cost for the same volume. Add a simple measure of serious errors caught before release. Account for setup, training and changing tool settings, especially if volume is low. Check whether the person reviewing remains capable of completing the task manually; that skill is part of the fallback. A positive result on one kind of note does not justify using the tool for confidential client advice or high-stakes decisions.
Common mistakes include treating fluent wording as evidence of accuracy, accepting an answer with invented references, overlooking provider settings and automating a messy process that should first be clarified. A task may be safe as an internal draft yet inappropriate for direct customer communication. When an output could affect safety, financial rights, employment or a regulated service, get specialist input and stronger controls. Rules differ by country and change over time. Do not endorse a paid tool solely because a demonstration looks fast.
Next step: list three repetitive tasks, mark the data sensitivity and likely consequence of an error, then choose one low-risk candidate with a human reviewer. Time ten comparable cases and record correction rate and full cost before deciding whether the tool deserves a place in everyday work.
Sources
- NIST: AI Risk Management Framework Core. Context, measurement, risk management and human oversight.
- UK Information Commissioner's Office: Generative AI and data protection consultation series. Questions about using generative AI with personal data.
- UK National Cyber Security Centre: AI and cyber security. Guidance on considering security from the beginning of an AI workflow.
