<oai_dc:dc xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd">
  <dc:creator>Sovrano, Francesco</dc:creator>
  <dc:creator>Dominici, Gabriele</dc:creator>
  <dc:creator>Langheinrich, Marc</dc:creator>
  <dc:date>2026</dc:date>
  <dc:description xmlns:ns0="xml" ns0:lang="en">A central goal of explainable AI is to express large language model (LLM) decision logic symbolically and ground it in internal mechanisms. Existing rule-extraction methods usually learn ungrounded symbolic surrogates, while mechanistic interpretability links behavior to neurons but often requires hand-crafted hypotheses and costly interventions. We introduce MechaRule, a pipeline that grounds rule extraction in LLM circuits by localizing sparse agonist activations whose ablation disrupts rule-related behavior. MechaRule rests on two findings. First, in a fixed baseline/flip regime, sparse agonist effects can exhibit overtopping: a few high-effect activations remain detectable within larger groups, dominate weaker ones, and flip many of the same examples. In such regimes, adaptive group testing with confidence-guided conservative pruning requires O(k log N/k +k) interventions over N candidates when k « N are agonists. Second, agonists are localized more reliably on data splits aligned with close-to-faithful rule behavior; spectral splits provide a rule-free fallback, whereas unfaithful splits degrade localization. Empirically, on arithmetic and jailbreaking, MechaRule recalls 97.0% of highest-effect agonists in matched brute-force validations at only 2.14% of exhaustive-ablation cost on average. Ablating the localized agonists eliminates 97.6-100.0% of eligible correct arithmetic answers and jailbreaks, and can correct arithmetic errors or induce jailbreaks by up to 72.8% and 32.5%.</dc:description>
  <dc:format>application/pdf</dc:format>
  <dc:identifier>https://localhost:5000/ark:/12658/srd1336637</dc:identifier>
  <dc:identifier>https://susi.usi.ch/global/documents/336637</dc:identifier>
  <dc:identifier>https://susi.usi.ch/documents/336637/files/3770855.3818091_compressed.pdf</dc:identifier>
  <dc:language>eng</dc:language>
  <dc:relation>info:eu-repo/semantics/altIdentifier/doi/10.1145/3770855.3818091</dc:relation>
  <dc:relation>info:eu-repo/semantics/altIdentifier/ark/12658/srd1336637</dc:relation>
  <dc:rights>info:eu-repo/semantics/openAccess</dc:rights>
  <dc:rights>CC BY</dc:rights>
  <dc:source>Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD ’26), August 09–13, 2026, Jeju Island, Republic of Korea. - New York : ACM - Association for Computing Machinery. - 2026, vol. 2, p. 4369-4380</dc:source>
  <dc:subject xmlns:ns1="xml" ns1:lang="en">Explainable AI</dc:subject>
  <dc:subject xmlns:ns2="xml" ns2:lang="en">Rule extraction</dc:subject>
  <dc:subject xmlns:ns3="xml" ns3:lang="en">Mechanistic interpretability</dc:subject>
  <dc:subject xmlns:ns4="xml" ns4:lang="en">Contrastive hierarchical ablation</dc:subject>
  <dc:subject xmlns:ns5="xml" ns5:lang="en">Neuron activation analysis</dc:subject>
  <dc:subject xmlns:ns6="xml" ns6:lang="en">Large Language Models</dc:subject>
  <dc:subject>info:eu-repo/classification/udc/004</dc:subject>
  <dc:title xmlns:ns7="xml" ns7:lang="en">Neuron-Anchored Rule Extraction for Large Language Models via Contrastive Hierarchical Ablation</dc:title>
  <dc:type>http://purl.org/coar/resource_type/c_5794</dc:type>
</oai_dc:dc>
