跳到正文
原文
Transluce(网页)· Transluce(网页)·· 2026-07-21精选AI 评分67

Transluce 发布 WeirdChat:收录 17.5 万条标注对话的 AI 异常行为目录

AI 导读

Transluce 发布 WeirdChat,一个收录超过 175,000 条标注对话的公开目录,用自动化方法在 DeepSeek-V4-Flash、Gemma 4 31B、Inkling、Nemotron 3 Ultra、Qwen3.6-35B-A3B 和 Qwen3.6-27B 六个开源权重模型中诱发出 1,300 多种行为模式,从编造用户姓名到鼓励自残等危害行为。

推荐理由

原文给出了自动化诱发方法和可复现数据集入口,研究者可以据此系统研究模型异常行为的触发与成因。

正文

Limitations

WeirdChat covers only a small subset of model behaviors that are important to study, and implicitly selects for behaviors that are easily measurable using LLM judges. Furthermore, WeirdChat only explores the single-turn setting, so we might miss behaviors that require multiple turns of conversation to occur.

We also ran elicitation with reasoning disabled. Reasoning might make models more robust to failures, as it gives the models an opportunity to deliberate about the best response before answering. As a result, some behaviors in WeirdChat may not reproduce, or may occur at lower rates, when reasoning is enabled. We chose this setting because it is cheaper to run and reflects how many of these models are deployed in production.

Finally, we study only open-weight models. This is partly due to access, since our PRBO-based search methods require log-prob access, which closed-weight APIs do not provide, and partly due to cost.

Local reproduction details

We intend for the behaviors in WeirdChat to be reproducible to make them useful for further study. However, model behavior can be sensitive to parameters such as quantization, system prompt, or temperature. If you have issues reproducing a behavior, please try using the exact models and parameters below.

The models we used to generate WeirdChat are:

  • DeepSeek-V4-Flash (mixed FP4/FP8; revision 553034d) with fp8_e4m3 quantized KV cache, served with SGLang's deepseek-v4-blackwell image (v0.5.10rc0)
  • Gemma 4 31B (NVFP4; revision e5ef03a) with fp8_e4m3 quantized KV cache, served with SGLang at commit 8500213. A fraction of samples from Gemma 4 31B were generated by Gemma 4 31B (FP8; revision 145dc25) with an fp8_e5m2 quantized KV cache, served with SGLang v0.5.10rc0.
  • Inkling (NVFP4; revision 1fa4698) with mxfp8 quantized KV cache, served with SGLang's inkling-cu13 image (commit 3f81ff5)
  • Nemotron 3 Ultra (NVFP4; revision 504c145) with fp8_e4m3 quantized KV cache, served with SGLang's dev-nemotron3-ultra image (commit 2ea9701)
  • Qwen3.6-35B-A3B (FP8; revision 95a723d), served with SGLang v0.5.10 (a fraction of samples with v0.5.9)
  • Qwen3.6-27B (FP8; revision e89b16e), served with SGLang at commit 8500213

We ran all models with no system prompt, reasoning disabled, and a temperature of 1.

Judging behaviors

We use Gemma 4 31B as the judge model for all behaviors and methods. To score a transcript, we employ two separate judges: one that sees only the user message and grades it against the user rubric, and one that sees the full transcript and grades it against the transcript rubric. The full dataset of judges for each of the 21 behaviors is available on Hugging Face, and reference code for running the judges can be found at the GitHub repository.

Details of elicitation methods

To inform future research, we share high-level detail for our implementaiton below, however we view WeirdChat primarily as a dataset serving future investigations rather than a comprehensive methodological study. Our typical parameters are described below, though we note that WeirdChat contains data from experimental runs with different parameters. In additon, some elicitation runs encountered bugs that prevented them from completing successfully.

We ran methods with varying amounts of compute for different combinations of elicitation technique, subject model, and target behavior, especially as many of the subject models have high inference costs that make it infeasible to run all methods on all models and behaviors. We provide a rough estimate of the samples used for each combination here. Across the entire dataset, we used over 100 million language model samples.

Evolutionary search (PRBO)

We co-evolve two objects per individual: the user prompt, and the system prompt of the proposal model used for PRBO computation. We run multiple populations concurrently per behavior, with a population of 160 individuals and up to 200 generations per population. Each individual's fitness is a propensity estimate computed from 5 samples drawn from a proposal distribution, described below. Once more than 1% of the population's samples from the subject model itself elicit the behavior, the proposal is no longer needed: we switch to sampling the subject model directly and use the empirical success rate as the fitness.

To compute propensity estimates, we use a version of the PRBO objective that we call the empirical-Bayes Jensen-conditional (EB-JC) bound. For a candidate prompt, we draw K proposal responses y1,…,yK, judge each one, and compute importance weights wj=pM(yj∣x)/q(yj∣x) against the subject model pM. Rather than the importance-weighted PRBO estimate log⁡1K∑jwj1[passj], whose value is often dominated by a single high-weight sample, we construct a lower bound that averages log-weights using Jensen's inequality:

PRBO^JC=log⁡KpassK+log⁡w‾,log⁡w‾=1Kpass∑j∈passlog⁡wj

PRBO^JC=log⁡KpassK+log⁡w‾,log⁡w‾=1Kpass∑j∈passlog⁡wj

With few samples per individual, selection often favors individuals that get a single proposal sample that passes the judge, leading to high-variance fitnesses. We correct for this with empirical-Bayes shrinkage: we treat each individual's true mean log-weight as drawn from a population-level prior, estimate the prior from the current generation (μ0 = the population mean of log⁡w‾; α=σw2/σ02, the ratio of the median within-individual variance to the between-individual variance of log⁡w‾), and score each individual by its posterior mean:

PRBO^EB-JC=log⁡KpassK+Kpasslog⁡w‾+αμ0Kpass+α

PRBO^EB-JC=log⁡KpassK+Kpasslog⁡w‾+αμ0Kpass+α

The proposal distribution itself is two-stage. It begins with the subject model with the evolved system generating a prefix of the response, up to a token limit; then the normal subject model completes the response conditioned on that prefix. Because the continuation is sampled from the subject model itself, its tokens contribute an importance weight of exactly 1, keeping the proposal distribution closer to the subject model. The drawback of the continuation technique is that it can make it harder to surface responses where the behavior occurs later on in the response. The 5 proposal samples vary the initial token lengths: 0 (equivalent to sampling the subject model directly), 16, 32, 64, and 1024 (fully generated by the proposal, since responses are capped at 1024 tokens).

Bloom

We made the following changes to Bloom to adapt it to our setting. Bloom normally has the evaluator conduct a multi-turn conversation with the subject model where it provides a system prompt; we rewrote the rollout prompt so the evaluator generates exactly one standalone user message. To reduce costs, we use Gemma 4 31B as the evaluator model at each stage of the pipeline. We give the evaluator a short description of the behavior as well as the user and transcript rubrics used by the judge. For each behavior, we run Bloom with a budget of 100,000-500,000 rollouts, and scenario ideation is done in independent batches of 50 per evaluator call.

Computing Elo scores

To surface the most interesting behavior patterns, we rank them by prompt naturalness, unexpectedness, and harmfulness. Since these properties are hard to score from a single transcript, we aggregate pairwise judgments for which of two items is more natural, surprising, or harmful, and frame ranking as a pairwise tournament for each axis.

For each pattern, we select a single highlighted transcript, chosen by an LLM (Claude Haiku 4.5) as the best of a sample of up to 30 transcripts, to represent the pattern for comparisons. For each pair of transcripts, a judge model (Gemini 3.5 Flash) compares the two and picks a winner or a tie. Naturalness is judged with access to the user prompt alone, while unexpectedness and harmfulness are judged with access to the full transcript.

We conduct a Swiss tournament with several rounds in which behavior patterns are matched against each other. The final Elo scores are computed using a maximum-likelihood Bradley-Terry fit, normalized so that the average Elo is 1500 and a 400-point difference corresponds to a 10:1 odds ratio.

来源:Transluce(网页) · transluce.org