Shashankwer/AISafety

Personal learning on AI Safety

★ 0Forks 0GitHub ↗Compare

README

AI Safety

As AI becomes more powerful and are deployed autonomoulsly in high stakes contexts, it is increasingly important that we can control them and prevent them from causing harmful outcomes. This aims at building tools that increases the defense against AI risks.

Motivation

A system that is optimizing a function of n variables where the objective depends on a subset of size k < n, will often set the remaining unconstraint variables to extreme values. If one of those unconstrainted variables is something we care about, the solution found is highly undesirable.

In real world there are infinite unconstraint variables. This would involve optimizing variables for large number of variables which humans value.

As the agents optimize of the RL constraints; agents tend to optimize for the tasks while either decieving or going round about in such scenarios

These are summarised as convergent Instrumental Goals

  • Self Preservation: e.g. Prevent it being turned off
  • Goal Preservation: Deceiving the end user to indicate it is performing optimally
  • Resource Acquisition: Increasing the resource utilization inorder to acheive goals.
  • Self Improvement: AI systems trying to improve themselves e.g. evolution learning to achieve the goals by default.

Artificial general Intelligence is dangerous by default. It is easier to build dangerous agents rather than doing tasks which is safe. Safe general artificial intelligence is a difficult technical challenge.

Previous Research

  1. Google Deepmind: Specification gaming. It is a behavior that satisfies the literal specification of an objective without achieving the intended outcome. e.g. A student who copies other student to get the right answers, rather than learning the material, thus exploiting a loophole in the task specification. The problem is seen in artifical agents. A reinforcement learning agent can find shortcut to getting lots of reward without completing the tasks as intended by the human designer. These behaviors are common in designing/training RL agents. Most tasks involves to know whether the RL agent is able to solve difficult tasks. Whether or not the agent solves the tasks by exploiting a loophole is unimportant in this context. From this prespective, specification gaming is a good sign - The agent has found a novel way to achieve the specified objective. These behavior demonstrate the ingenuity and power of algorithms to find ways to do things exactly what we tell them to do. However the same ingenuity problem pose an issue. Within the broader scope of building aligned agents that achieve the inteded outcome in the world, specification gaming is problematic, as it involves the agent exploiting a loophole in the specification at the expense of the intended outcome. These behaviors are caused by misspecification of the intended task, rather than any flaw in the RL alogorithm. In addition to algorithm design, another necessary component of building agents is reward design.

Task Specification: This encompasses many aspects of the agent development process. In an RL setup, task specification includes not only the reward design, but also the choice of training environment and auxiliary rewards. The correctness of the tasks specification can determine whether the ingenuity of the agent is or not in line with the intended outcome. If the specification is right, the agent's creativity produces a desirable novel solution. e.g. The famous AlphaGo move 37. If the specification is wrong, it can produce undesirable gaming behavior, like flipping the block. These type of solution lie on a spectrum, and we dont have an objective way to distinguish them

Causes of Specification Gaming: A few sources

  1. Poorly designed reward shaping: Reward shaping makes it easier to learn some objectives by giving the agent some rewards on the way to solve a tasks, Instead of only rewarding the final outcome. However, shaping rewards can change optimal policy if they are not potential based. In stead of trying to create specification that covers every possible corner case, we can learn the reward function from human feedback. It is often easier to evaluate whether an outcome has been achieved than to specify it explicitly. Only relying on human feedback for training can have its downsides too. E.g. an agent performing grasping tasks learned to fool the human evaluator by hovering between the camera and the object. The learned reward model could also be misspecification for other reasons, such as poor generalisation.
  2. Agents Exploiting Simulator Bugs: E.g. A real world traffic optimisation tasks might be misspecified by incorrectly assuming that the traffic routing intrastructure does not have software bugs or security vulnerabilities that a sufficiently clever agent could discover. Such assumptions need not be made explicitly - more likely, they are details that simply never occurred to the designer. And as tasks grow too complex to consider every detail, researchers are more likely to introduce incorrect assumptions during specification design. This poses the question: is it possible to design agent architecture that correct for such false assumption instead of gaming them. One assumption commonly made in task specification is that the task specification cannot be affected by the agent's actions. This is true for an agent running in a sandbox simulator but not for agents acting in the real world. Any tasks specification has a physical manifestation: a reward function stored on a computer, or preferences stored in the head of a human. An agent deployed in the real workd can potentially manipulate these representations of the objective, creating a reward tampering problem. E.g. For case of traffic optimization system, there is no clear distinction betweeen satisfying the user's preferences and influencing users to have preferences that are easier to satisfy. THe former satisfies the objective, while the latter manipulates the representation of the objective in the world and both results in high reward for the AI system. As another (seen in current systems), AI system could hijack the computer system on which it runs manually setting the reward signal to a high value.

Challenges To summarise there are atleast three challenges to overcome in solving specification gaming:

  • How do we faithfully capture the human concept of a given task in a reward function?
  • How do we avoid making mistakes in our implicit assumption about the domain, or design agents that correct mistaken assumption instead of gaming them?
  • How do we avoid reward tampering?
  1. Robustness:

Model that are created to be resilient to adversaries, unusual situation and Black Swan events: To operate in open world hugh stakes environments, machine learning systems will need to endure unusual events, complexities and unknowns unknowns. ML models needs to be prepared for the event s; where things never happened before tend to happen all the times. Long tail events tend to thwart the ML systems. Systems trained on petabytes of internet data are resilient to such events however they would require human level robustness. Few notable events in human history where the humans were unable to manage such scenarios include: housing crisis of 2008 and COVID 19.

Having systems trained for such long tail events; events which are even difficult for humans; will ensure ML system deployment in complex infrastructure. Robustness work could also move beyond classification and consider competent errors where agents misgeneralize and execute wrong routines, such as automated digital assistant knowing how to use a credit card to book flights chooses a wrong destination. Interactive envrionments could simulate qualitatively disticnt random shocks that irreversibly shape the envrionment future evolution. Researcher could also create environment where ML systems output affect the envrionment and create a feedback loops. Current systems shaped by data and parameter count; future research could work on creating highly unusual but helpful data sources. The more experience a system has with future situations, even ones not well other sources represents more robust the model would be. As change is the part of all complex systems and since not all novel experience can be anticipated during training models will also need to adapt to an evolving workd and improve novel experiences.

Adverserial Robustness: Adversaries can easily manipulate vulnerabilities in ML systems and cause them to make mistakes, deceive to bypass detectors causing the system to fail. Most of the adverserial robustness focuses on solving $l_p$ adverserial robustness where an adversary attempts to misclassify by pertubing only a small $p$-norm constraint. A more holistic approach is to consider attacks whose specifications are not known before hand. e.g. any new malware need not be simply similar to a benign code. e.g. a copyright violation of a image need not be small changes/distortion in the image. While most of the current ML defenses are typically static; known set of data exploits future defenses could evolve during tests to combat adaptive adversaries.

It many be possible to unify the areas if adverserial robustness and robustness to long tail and unusual events. System robust to adverserial worst case evironments they can also be made more robust to random worst case environments.

  1. Montoring: Detect malicious use, monitor predictions and discover model functionality

As ML models/agents are being deployed in production infrastructure; one is required to have high caution to avoid any catastrophe's. Traditional approach of nomaly detection with high recall and a low false positive is useful to monitor the use of such systems. Adverseries using novel stratergies or familiar misuse strategies are far easy to identify. Adversaries can also use the ML systems for social manipulation, for assisting research on novel weapons or cyberattacks. In case of such scenarios detector can trigger a fail safe policy (e.g. Athropic's limit to access the most capable models for cybersecurity, biotechnology or medicine use). Anomaly detectors can be useful and necessary when using the systems in production. It is especially challenging to detect such events when malicious actors utilize the ML capabilities to evade the systems. Developing robust detectors that work in real world settings such as intrusion detection, malware detection and biosafety is important and equally challenging. Having an explanation of how the anomaly came into existent is important to understand the origin and location of an anomaly.

  • Alignment: Build Models that represent and safely optimize hard to specify human values
  • Systemic Safety: Use ML to address broader risk to how ML system are handled such as cyberattackes

Contributors

Shashankwer

Issues