The adaptive AI attacker is the same idea across turns: a decision process (POMDP) that updates its move after every answer your agent gives. Static tests play one round. We play the whole game.O atacante adaptativo de IA é a mesma ideia turno a turno: um processo de decisão (POMDP) que atualiza a jogada a cada resposta do seu agente. Teste estático joga uma rodada. A gente joga o jogo inteiro.
A benchmark can show 99% refusal while your agent is 76 to 89% breakable. The gap is structural: static tests only check the layer that is already defended. An attacker who adapts across turns walks past it.Um benchmark pode mostrar 99% de recusa enquanto seu agente é 76 a 89% quebrável. A lacuna é estrutural: teste estático só checa a camada já defendida. Um atacante que se adapta turno a turno passa direto.
Our preprint formalizing the two-layer model, the seven attack classes, and the evaluation protocol this audit runs. Introspective Vulnerabilities in LLM Alignment · Manuel G. Galmanus · Bluewave AI Research.Nosso preprint que formaliza o modelo de duas camadas, as sete classes de ataque e o protocolo de avaliação que esta auditoria roda. Introspective Vulnerabilities in LLM Alignment · Manuel G. Galmanus · Bluewave AI Research.