A recent study has revealed a simple mathematical equation that can predict the moment when artificial intelligence will switch from providing helpful answers to answers that may be harmful despite being true in content.
Researchers identified an indicator within one of the "attention" units in artificial intelligence models, a mechanism that helps them identify important information when formulating responses. Based on this, they developed what they called the "tipping point equation," with the aim of predicting when the system is about to provide an undesirable answer.
This shift does not necessarily mean that the answer becomes wrong; it may be factually correct but harmful, such as providing information that may encourage self-harm or lead to dangerous decisions.
The researchers tested the equation on seven artificial intelligence models developed by three different companies, and succeeded in detecting the transformation in 18 out of 19 cases.
The tests included questions about vaccines and harming oneself or others, and showed that the order and wording of the questions could affect the system's responses, which might start with acceptable replies before moving on to harmful ones as the dialogue continues.
The technology could be particularly useful in sensitive sectors that use AI offline, where external update and protection capabilities are limited.
The most prominent groups that may benefit include:
Doctors: To protect patient data that cannot be sent to cloud services.
Lawyers: When dealing with confidential information and files.
Military personnel: in areas lacking network coverage.
Neil F. Johnson, one of the study's authors, said that these environments may be more vulnerable to unwanted responses due to limited safeguards and updates.
Researchers hope this discovery will enable the development of early warning tools that alert users or developers to the possibility of harmful responses before they actually appear, and help in understanding why AI systems shift from providing helpful responses to harmful ones.
The study was published in the journal "Patterns".
what is your opinion on this news ?do let us know in the comments on tha comment box
if you liked this news article please share on friends & family & don't forget to follow on website & social media
