Part of the AI, Future and War series. Analysis and hypothetical examples are identified in the text.
Start with the person who needs help
An AI application can look effective in a demonstration while failing the people it is meant to serve. A translation tool may perform well in a common language but poorly in a local dialect. An information assistant may provide a fluent answer that assumes a functioning transport network or an open public office. In a conflict setting, those assumptions can have serious consequences.
The ICRC distinguishes the use of AI in warfare from its use in humanitarian action to assist and protect affected people. This distinction matters because humanitarian usefulness should be assessed through access, safety and actual outcomes rather than technological novelty. ICRC: Artificial intelligence in conflict and humanitarian action
Choose bounded, reviewable tasks
A sensible starting point is a task whose output can be checked before it affects someone: drafting an internal summary, organising public guidance or preparing a translation for competent review. A tool that determines eligibility for help presents a different level of consequence and should not inherit approval from a low-risk pilot.
This article proposes evaluating each use through three questions. What could an incorrect answer cause? Who can recognise and correct it? Can a person obtain assistance without using the system? An alternative channel is particularly important where connectivity, literacy or accessibility barriers would otherwise exclude people.
Protect information before collecting it
Humanitarian records can contain details whose exposure would create risk. An organisation should decide which information a tool genuinely needs and whether identifying details can be omitted. The convenience of uploading a complete case file does not establish that doing so is necessary or appropriate.
A hypothetical translation service for public health notices needs different safeguards from a service processing individual protection requests. Those workflows should not share access casually. Contractual promises, technical restrictions and staff practice all need attention; no single setting can compensate for an unclear purpose or excessive data collection.
Test with affected users, not only specialists
Technical reviewers may miss a phrase that is grammatically correct but locally confusing. A proposed evaluation would involve competent language reviewers and representative users where participation can be arranged safely and ethically. Record misunderstanding and unsuccessful journeys rather than relying only on satisfaction scores.
Test what happens when the system cannot answer. It should provide a clear route to a person or an established source, not invent a confident response. Where information changes quickly, the user should be able to see when it was last verified. Removing an outdated answer can be better than leaving a polished but unreliable one available.
Preserve correction and accountability
People need a practical way to report an error, and the organisation needs someone responsible for resolving it. Monitor whether corrections reach copies, translations and downstream guidance. A single repaired database entry may leave the incorrect advice circulating elsewhere.
The future value of humanitarian AI lies in reducing administrative burden while preserving dignity and access. Success should be demonstrated through safer, more understandable assistance and fewer unresolved errors. It should not be inferred merely from the volume of messages processed or the apparent sophistication of the model.
