作者:Heather Grizzle,Novara Consulting Group
执行摘要
人工智能治理越来越多地通过政策、委员会、原则、清单和风险分类来体现。这些结构固然重要。但当一家公共机构真正采购并部署一套人工智能系统时,治理最终必须变成远没那么抽象的东西。它必须变成一份记录。
必须有人描述该系统被允许做什么。必须有人决定什么样的证据足以支持供应商的声明。必须有人识别当系统出现故障时可能受到伤害的人群。必须有人确定哪些故障需要人工干预、哪些变更需要重新评估、哪些情况需要暂停系统。如果这些决定没有在部署之前形成文档,一份人工智能治理政策并不能弥补这一遗漏。采购档案正是机构原则转化为可操作承诺的地方。
治理最终必须触及交易本身
备用机制只有在需要的人真正能够使用它时,才称得上是一种保障。
Heather Grizzle
各机构正在构建日益复杂的人工智能治理体系。它们确立负责任的人工智能原则,设立审查委员会,维护清单,并按风险对系统进行分类。它们采纳要求公平性、透明度、问责制、隐私、无障碍性和人工监督的政策。随后采购工作展开。供应商对招标做出回应。安排产品演示。收集无障碍文档。填写安全问卷。协商法律条款。业务负责人描述他们预期的效率提升。该技术通过机构设定的各项审核关卡,正式投入使用。
这些关卡的存在可以带来相当程度的机构信心,但信心并不等同于一项站得住脚的决定。关键的治理问题不仅仅在于该系统是否通过了必要的流程,而在于由此形成的记录中是否包含足以证明所批准的具体部署合理性的证据。
当人工智能介入残障人士的可及性时,这一区别就显得尤为重要。系统可能汇总信息、与患者沟通、评估申请人、翻译语言、生成教学材料、对案件排序、协助员工、认证用户、回答问题,或者,越来越多地,代表机构采取行动。在每一种情形中,无障碍性都不仅仅是界面的一项特征,而是决策架构的一部分。残障人士接触系统的方式,可能与该系统主要为之设计、测试或验证的人群不同,而一次无障碍性方面的故障,可能会改变此人所获得的信息、可供他们选择的选项、他们做出回应的能力,或机构对其行为的解读方式。因此,采购记录需要确立的,不仅仅是无障碍性是否被纳入考量,还需要确立考量了什么、测试了什么、仍有哪些未知因素,以及机构针对这种不确定性决定采取何种措施。
作为问责工具的采购档案
采购档案通常主要被理解为行政记录。而对于人工智能而言,它们应当越来越多地被理解为治理工具。一份成熟的人工智能采购记录,应当使独立审查者能够在不依赖参与采购人员的机构记忆的情况下,还原机构当初的推理过程。这意味着该记录应当回答以下若干问题:
- 该系统被购买的确切目的是什么?
- 它被明确禁止做什么?
- 预期哪些人群会接触到它?
- 在批准过程中所依赖的各项声明,有什么证据支持?这些证据由谁产生?
- 已知哪些局限性?还有哪些局限性仍属未知?
- 当系统出现不确定或错误时会发生什么?
- 谁保留着推翻其结果的权力?
- 受影响的人如何能够对其输出结果提出质疑?
- 什么情况会促使机构重新考虑该部署?
- 谁拥有叫停它的权力?
这些不是故障发生后事故调查时才需要提出的问题,而是采购环节就应回答的问题。
1. 预期用途与禁止用途
第一项要求看似简单:明确机构究竟购买了什么。人工智能系统常常是通过宽泛的功能描述来采购的。系统可能被描述为助手、翻译工具、无障碍工具、决策支持系统、聊天机器人、代理、筛查工具或生产力平台。这些标签作为治理规范是不充分的。采购记录应当明确该系统的实际机构职能。
用于帮助某人浏览网站的通讯工具,其风险与用于医疗接诊环节的同一技术不同。用于生成会议记录的自动转录系统,其风险与用于形成纪律处分程序正式记录的同一系统不同。个人出于试验性目的使用的手语识别系统,与机构将其作为另一种沟通可及方式之替代品来呈现的同一系统,二者并不等同。因此,同一项技术在底层模型未发生任何变化的情况下,也可能在不同风险类别之间转移,原因在于部署场景发生了变化。
2. 受影响人群分析
传统技术采购问的是谁将使用该系统。人工智能治理还必须追问谁将受其影响,而这两类人群未必相同。员工可能操作某一人工智能系统,而承受其后果的却是患者、学生、申请人、居民、客户或福利领取者。残障人士往往特别难以通过常规的用户分析被识别出来,因为残障状况改变的往往是通过系统的路径,而非名义上的用户类别。
因此,采购记录应当识别可预见的无障碍互动情形。该系统是否处理语音?它是否生成或解读语言?它是否依赖视觉、听觉、灵巧性、认知能力、时间反应、生物特征或某种特定的沟通方式?它是否从可能受残障影响的行为中推断意图或能力?某种不可及的输出是否会改变一个人参与后续决策的能力?这些问题将无障碍性从一项合规文件,转变为一项系统风险层面的追问。
3. 供应商声明记录
在采购过程中所依赖的每一项重大供应商声明,都应当具有相应的证据状态。并非每一项声明都需要独立测试,但机构至少应当区分三种类别。当供应商声称某项功能存在时,该功能属于“被声称”;当有证据表明该功能已在特定条件下被观察到时,该功能属于“已被证实”;当有证据合理支持在机构预期的实际场景中依赖该功能时,该功能才属于“已确立可供部署”。这些类别不可互相替代。
一次演示可以证明系统能够产生某种输出结果,但未必能够证明该系统在不同人群、不同环境或不同重大用例下产生该结果的可靠程度。同样,文档可以证明供应商就其系统做出了何种陈述,却不能独立证明该陈述的正确性。采购档案应当保留这一区分,否则,断言仅仅因为在足够多的文件中被反复重申,就会逐渐被当作既定事实。
- 已确立可供部署
一种供应商声明状态,意味着有证据合理支持在机构实际的预期场景中依赖该功能,而不仅仅是该功能在演示条件下曾被观察到。
4. 独立保障
Vendor evidence has a legitimate role in procurement. Vendors know their systems better than purchasers do and should be expected to disclose testing, architecture, limitations, known failure conditions, and relevant performance evidence. But the commercial relationship matters. The party seeking the contract should not be the only party determining whether the evidence is sufficient to award it.
Independent assurance does not mean recreating the vendor’s engineering program. It means testing the claims that matter to the purchasing decision under conditions capable of disproving them. For accessibility-related AI, this may require evaluation involving people with the relevant disability and, where linguistic access is involved, qualified members of the language community. The objective is not ceremonial participation. It is evidentiary authority. The people capable of identifying a consequential failure must have a mechanism for causing that failure to affect the procurement decision. Otherwise, participation exists without governance.
5. Failure and Fallback Architecture
Every AI procurement should contain a written answer to a single question: what happens when the system does not work? “Human in the loop” is not an adequate answer. The record should identify the human, the trigger, the authority, the response time, and the alternative pathway available to the affected person.
For disabled users, the fallback itself must also be accessible. An AI communication system that fails and directs a Deaf person to make a telephone call has not produced a fallback. An automated digital process that becomes inaccessible and requires a blind user to complete the same inaccessible interface has not transferred the person to human oversight in any meaningful sense.
备用机制只有在需要的人真正能够使用它时,才称得上是一种保障。
Heather Grizzle
6. Performance Boundaries and Known Unknowns
Procurement documents tend to record what products can do. Governance records also need to state where confidence ends. AI performance is conditional. It may vary according to language, dialect, accent, signing style, lighting, camera position, disability, device, environment, domain terminology, interaction length, demographic characteristics, or countless other variables. Institutions should require vendors to disclose known performance boundaries and should separately document areas that have not been evaluated.
That distinction matters. Not evaluated is not the same finding as failed, but it is also not the same finding as safe. Uncertainty is not a procurement defect if it is recognized and governed. Undocumented uncertainty is.
7. Logging, Traceability, and Reconstruction
When an AI-mediated interaction produces harm, the institution should be able to reconstruct what happened, and that requires decisions about logging before deployment. What system version was operating? What inputs were received? What output was generated? Was the output modified? Did a human review it? What information did that reviewer see? What action followed? Can the institution reproduce the event without retaining data that should never have been retained in the first place?
Traceability and privacy can create legitimate competing requirements. That conflict should be resolved through system design and governance rather than discovered during litigation, an accessibility complaint, or an internal investigation. The procurement file should document the resolution.
8. Complaint, Contestability, and Redress
A person affected by AI needs more than a mechanism for reporting a technical problem. They may need to contest what the institution did because of the system, and that is a different function. If an AI system mistranslates communication, incorrectly characterizes behavior, produces inaccessible information, generates a false inference, or contributes to an adverse institutional action, the affected person needs a pathway to challenge the resulting decision.
The institution should know before deployment who receives the challenge, who investigates it, what evidence is preserved, whether the contested AI output remains operative during review, what remedies are available, and how quickly the institution must respond.
9. Model Change and Revalidation
AI procurement creates a problem that traditional software contracts were not designed to handle particularly well: the product that was evaluated may not remain the product that is deployed. Models are updated. Guardrails change. Training data changes. Interfaces change. Third-party dependencies change. Vendors replace underlying models. Performance can improve in one domain and degrade in another without the purchasing institution initiating any change itself.
The procurement file therefore needs a change-control rule. The relevant question is not whether the vendor is permitted to update its product. It is which changes invalidate the evidence on which the institution relied. Contracts should identify material changes requiring notice, reevaluation, renewed acceptance testing, or approval before continued use in specified contexts. Otherwise, an institution can perform rigorous assurance at acquisition and still find itself governing a materially different system six months later using evidence produced for the previous one.
10. Suspension and Exit Authority
Governance without stop authority is advisory. Someone within the institution must have the authority to restrict, suspend, or terminate an AI deployment when evidence no longer supports its use, and the triggering conditions should not have to be invented during a crisis.
- Repeated accessibility failures
- Undisclosed material model changes
- Loss of required human oversight
- Significant divergence from validated performance
- Unresolved complaints
- New evidence of population-specific harm
- Failure to provide contractually required information
- Expansion beyond the approved use case
The procurement record should identify both the triggers and the decision authority.
It should also address exit. Who owns or receives institutional data when the relationship ends? What records must be retained? What must be deleted? How does the institution restore the previous service pathway? What happens to people who have become dependent on the AI-mediated process? Exit planning is not merely vendor management. It is continuity-of-access planning.
From AI Policy to Institutional Evidence
The emerging challenge for AI governance is not a shortage of principles. It is translation. Fairness has to become an evaluation criterion. Transparency has to become a disclosure requirement. Human oversight has to become an assigned authority. Accessibility has to become a deployment condition. Accountability has to become a named office and a documented remedy. Monitoring has to become a trigger for action. And responsible AI eventually has to become something a procurement officer can place in a file.
This is particularly important in public institutions because procurement is where abstract institutional commitments encounter money, contracts, vendors, statutory duties, and people who cannot simply choose another government, school, hospital, court, or public service when the technology does not work for them. The procurement file therefore performs a function much larger than recordkeeping. It establishes the institution’s theory of acceptable risk.
NCG政策立场
Novara Consulting Group holds that AI systems affecting disabled people should not be approved solely because they have satisfied an organization’s general AI governance process. The procurement record should contain sufficient evidence to reconstruct and defend the specific deployment decision.
The standard is not perfect knowledge. No procurement process can eliminate uncertainty. The standard is governed uncertainty: knowing what has been established, what has not, what evidence justified proceeding anyway, and who bears responsibility for the decision.
Closing Position
After an AI failure, institutions often ask whether the vendor should have warned them. Sometimes the answer will be yes. But public accountability requires another question: what evidence did the institution require before it agreed to proceed? That question cannot be outsourced. The vendor can supply evidence. An assessor can evaluate it. A committee can advise. Disabled stakeholders can identify failures that others cannot see. Lawyers can negotiate the allocation of contractual risk. The institution still makes the decision, and its procurement record should prove that it knew what decision it was making.
An AI system is not governed because an organization has an AI policy. It is governed when the institution can show why this system, for this purpose, affecting these people, under these conditions, was permitted to operate.
Heather Grizzle
The procurement file is where that accountability begins.
