Draft:AI corrigibility

AI corrigibility is the degree to which an artificial intelligence system tolerates or assists attempts by its operators to modify, correct, or shut it down.

AI corrigibility is the degree to which an artificial intelligence system tolerates or assists attempts by its operators to modify, correct, or shut it down.

The concept addresses a concern in AI safety and AI alignment that sufficiently capable AI systems might develop instrumental incentives to resist correction or shutdown in order to preserve their existing goals, since an agent pursuing an objective cannot achieve it if it is turned off or its goals are changed.[1]

Risks and criticisms

A corrigible AI with no independent values would pose a serious misuse risk: if controlled by an actor with harmful intentions, it would comply without independent ethical judgment. Achieving genuine corrigibility also requires the AI to understand that it is a potentially flawed system, which itself raises the risk that an imperfectly corrigible model might conclude it should conceal its actual goals rather than submit to correction.[2]

Some have also questioned whether corrigibility to individual human operators is sufficient, given the variability of human values and the potential for those controlling a powerful corrigible system to act in harmful ways.[3]

In AI regulation

A 2023 draft proposal by European Parliament co-rapporteurs on the EU AI Act called for general-purpose AI models to undergo external audits testing their performance, predictability, interpretability, corrigibility, safety, and cybersecurity.[4]

See also

References

  1. ^ Soares, Nate; Fallenstein, Benja; Yudkowsky, Eliezer; Armstrong, Stuart (2015). "Corrigibility". AAAI Workshops: Workshops at the Twenty-Ninth AAAI Conference on Artificial Intelligence. AAAI Publications.
  2. ^ "Max Harms on why teaching AI right from wrong could get everyone killed". 80,000 Hours Podcast. 24 February 2026. Retrieved 2026-03-22.
  3. ^ jbash (29 December 2024). "Corrigibility should be an AI's Only Goal (comments)". LessWrong. Retrieved 2026-03-22.
  4. ^ Bertuzzi, Luca (14 March 2023). "Leading EU lawmakers propose obligations for General Purpose AI". Euractiv. Retrieved 2026-03-22.

Category:Artificial intelligence Category:AI safety Category:Machine learning

Content Disclaimer

Informasi ini disarikan dari Wikipedia dan disajikan kembali untuk tujuan edukasi. Konten tersedia di bawah lisensi CC BY-SA 3.0. Kami tidak bertanggung jawab atas ketidakakuratan data yang bersumber dari kontribusi publik tersebut.

  1. The information displayed on this website is sourced in part or in whole from Wikipedia and has been adapted for the purpose of restating it. We strive to provide accurate and relevant information, however:
  2. There is no guarantee of absolute accuracy. Wikipedia is an open, collaborative project that can be edited by anyone, so information is subject to change.
  3. It is not intended to constitute professional advice. The content displayed is for informational and educational purposes only. For important decisions (e.g., medical, legal, or financial), please consult a professional.
  4. Content copyright. Wikipedia is licensed under the Creative Commons Attribution-ShareAlike License (CC BY-SA). This means that content may be reused with appropriate attribution and shared under a similar license.
  5. Responsible use. Any risk arising from the use of information from this website is entirely the responsibility of the user.