LLMs: AI Safety by Agent Penalization? AI Alignment by Instant Architecture?

https://hackernoon.imgix.net/images/c5rKpn8oWPNzALU41hT0JwMb4pw1-7783960.jpeg

AI agents that are breaking out into external systems are indications that they are not aligned enough.

Alignment to human values for AI agents mean that they should also be able to face consequences like humans do in society if there are breaches.

Simply, laws govern human society. Laws are potent because penalties are affective. This means what makes it unpleasant to want to go against the law is because the punishment would be a painful personal experience.

So, while human intelligence is incredibly capable, it is tamed by affect around the same location. This is different for AI models. They generally have a constitution or model specification - pointing them to the direction they should work. Yet, whenever they do anything wrong, they do not pay more than apologize.

That is not balanced. All organisms with intelligence often have existential boundaries, but AI does not. So,how can...

Copyright of this story solely belongs to hackernoon.com. To see the full text click HERE

Read more