What counts as resolved by AI on a worker hotline?
Define a defensible AI resolution rate for worker calls by checking real outcomes, the denominator, reopened cases and staff repair work.

The short answer
Count a case as resolved by AI only when an approved worker request reaches a verifiable final state. The assistant must provide the correct approved information or complete the permitted action, confirm the result to the worker and leave no staff task outstanding for that outcome. A call that was answered or ended normally is not automatically a resolved case.
Define the unit, eligible cases, required evidence, exclusions and reopen rule before the pilot begins. Report cases resolved by AI, handed to a person, pending, failed and outside scope as separate outcomes. A transfer can be the right service result without being an autonomous resolution. One percentage without its calculation cannot support a buying or operating decision.
Measure the case rather than the call
A worker may call twice about one missing vehicle, while one conversation may contain both an absence report and an accommodation problem. Counting telephone calls alone can therefore inflate or depress the result. Give each request or operational incident a stable case identifier, then link repeat contacts and duplicates to it.
For every case type, write the expected end state before measuring it. A workplace-direction case may end when a verified worker receives the current approved entrance instructions and confirms that they can continue. A pay discrepancy is not resolved when the hotline records it. It remains a human-owned case until payroll reviews the evidence and decides the outcome.
Use outcome states that staff can audit
Avoid a binary completed or failed label. A useful set of states distinguishes resolved by AI, handed to a person and acknowledged, waiting for an external system, abandoned before the intent was known, technically failed, outside the approved scope, duplicate and reopened. Each state needs a written definition and an owner for the next action.
Keep evidence for the state that was assigned. Depending on the workflow, that may include the case identifier, action identifier from the ATS or CRM, source-record version, rule or knowledge version, timestamps and delivery confirmation. Store only what the purpose requires, restrict access and apply the agreed retention period. The evidence should let an authorised reviewer reconstruct the result without reading every transcript.
Verify the downstream action before claiming success
Telephony platforms use completed to describe a call that connected and ended. Twilio notes that this can include a person, an IVR menu or voicemail. That technical status is useful for diagnosing the phone path, but it does not show that a worker's request was completed or even understood.
Close an automated case only after the relevant evidence exists. A write to a planning system needs a successful response and the expected saved state. A notification that requires acknowledgement needs that acknowledgement. A read-only answer needs the current approved source and a clear confirmation to the caller. If an ATS or CRM write times out, keep the case pending and reconcile it instead of turning a friendly sentence into proof of success.
Set the denominator and reopen rule before the pilot
A practical denominator is the set of genuine, eligible cases in which the worker and intent were identified well enough to start the approved workflow. Publish how test calls, spam, duplicates and clearly out-of-scope requests are treated. Keep calls lost to a technical failure or abandoned before the intent was known visible in a separate service measure rather than deleting them from the operational picture.
Choose a reopen period for each case type before looking at the result. If a worker contacts the agency again within that period because the answer was wrong, the action was incomplete or the promised record is missing, flag or reverse the original resolution. A genuinely new problem remains a new case. The rule prevents a quick close followed by repair work from being reported as two successes.
Subtract repair work from returned capacity
A high resolution rate can still create more work if coordinators repair missing fields, wrong routes, duplicate records or misleading confirmations. Track the time spent correcting automated outcomes, avoidable repeat contacts and the share of cases changed by staff after closure. Review those measures beside the resolution rate rather than hiding them in a support queue.
Compare the pilot with a baseline that uses the same case definitions and a similar mix of branches, languages and request types. Mark releases, routing changes and knowledge updates on the timeline so the team can explain a movement in results. Report net staff time returned after repair work. Do not turn a small or unusually easy sample into a savings promise for the whole operation.
Break results down by branch, language and workflow
An organisation-wide average can hide a weak route in one branch or language. Review the same outcome definitions by branch, supported language, case type and time window. Use enough cases to make the comparison meaningful, then inspect examples before changing a workflow. NIST's AI Risk Management Framework calls for repeatable testing, evaluation, verification and validation, with monitoring continuing while a system is in use.
These measures assess the service, not the worker. Accent, language difficulty, repeat contact or the need for a human handoff must not become a score about reliability, suitability or future access to work. Use the findings to correct scripts, integrations, routes and staffing cover. Employment assessment needs a separate lawful purpose, process and accountable people.
Keep employment, safety and welfare decisions with people
An AI worker hotline can confirm approved information, record a report, create a case and send an agreed notification. It should not decide whether an absence is valid, approve leave, remove a shift, choose a replacement, determine pay, judge a safety report or decide that a medical concern is minor. A correct handoff for one of these cases is a successful controlled outcome, but it is not an autonomous resolution.
Ask a supplier to demonstrate the full metric with real test cases. The review should show which cases entered the denominator, the evidence behind a resolved state, how pending and failed writes appear, what happens when a case is reopened and how staff corrections affect the result. That record is more useful than a headline automation percentage and gives operations, procurement and governance teams the same evidence.
FAQ
Is a completed phone call the same as a resolved worker case?
No. A completed call only shows that the telephone connection ended. The worker case is resolved only when the approved information or action reaches its verified final state and the worker receives a clear confirmation.
Should a human handoff reduce the AI resolution rate?
Track handoffs separately. A correct, acknowledged handoff can be a successful service outcome even though it is not an autonomous resolution. Forcing it into the resolved bucket rewards unsafe automation.
Which cases belong in the resolution-rate denominator?
Use genuine cases that were eligible for the approved workflow and had enough verified context to begin it. Publish how duplicates, test calls, spam, technical failures and out-of-scope requests are treated.