Zürcher Nachrichten - AI systems are already deceiving us -- and that's a problem, experts warn

EUR -
AED 4.17258
AFN 73.27737
ALL 91.859328
AMD 413.248933
ANG 2.034159
AOA 1041.867426
ARS 1732.627266
AUD 1.62491
AWG 2.046526
AZN 1.931008
BAM 1.955085
BBD 2.289773
BDT 139.978829
BGN 1.912675
BHD 0.42854
BIF 3406.48471
BMD 1.13617
BND 1.453182
BOB 13.920911
BRL 5.934891
BSD 1.136889
BTN 109.061652
BWP 15.544386
BYN 3.441342
BYR 22268.925387
BZD 2.286474
CAD 1.611531
CDF 2641.593994
CHF 0.946196
CLF 0.027846
CLP 1099.527859
CNY 7.624209
CNH 7.620926
COP 3829.368954
CRC 516.313529
CUC 1.13617
CUP 27.285146
CVE 110.224706
CZK 24.400777
DJF 202.446885
DKK 7.474667
DOP 67.685555
DZD 152.060918
EGP 59.158759
ERN 17.042545
ETB 185.266589
FJD 2.547009
FKP 0.857057
GBP 0.85803
GEL 2.948374
GGP 0.857057
GHS 13.227165
GIP 0.857057
GMD 84.076294
GNF 9998.169078
GTQ 8.682788
GYD 237.88304
HKD 8.912961
HNL 30.573572
HRK 7.532923
HTG 148.781028
HUF 368.000798
IDR 20478.321999
ILS 3.49204
IMP 0.857057
INR 109.111429
IQD 1488.950343
IRR 1562062.860277
ISK 136.988007
JEP 0.857057
JMD 179.967875
JOD 0.805566
JPY 178.684831
KES 147.473631
KGS 99.356336
KHR 4614.353793
KMF 491.961762
KPW 1022.553058
KRW 1540.71482
KWD 0.350847
KYD 0.947458
KZT 499.959435
LAK 25508.675066
LBP 101806.883545
LKR 376.302439
LRD 195.537659
LSL 18.682271
LTL 3.354814
LVL 0.687257
LYD 7.272374
MAD 10.948098
MDL 20.099652
MGA 4987.19883
MKD 61.552499
MMK 2385.242324
MNT 4087.054422
MOP 9.186242
MRU 45.543752
MUR 54.092851
MVR 17.56497
MWK 1971.388016
MXN 20.445712
MYR 4.637394
MZN 72.597398
NAD 18.682353
NGN 1501.982302
NIO 41.640411
NOK 10.849744
NPR 174.499978
NZD 2.007561
OMR 0.436864
PAB 1.136884
PEN 3.862952
PGK 5.049422
PHP 71.016322
PKR 314.775873
PLN 4.37277
PYG 6677.588332
QAR 4.141388
RON 5.278764
RSD 117.561742
RUB 96.011039
RWF 1678.527304
SAR 4.267355
SBD 9.119124
SCR 15.789353
SDG 683.412714
SEK 11.337593
SGD 1.452667
SHP 0.857131
SLE 27.914413
SLL 23824.900515
SOS 649.322841
SRD 42.817674
STD 23516.418098
STN 24.654882
SVC 9.947364
SYP 14772.478242
SZL 18.677255
THB 38.188943
TJS 10.48762
TMT 3.976594
TND 3.366669
TOP 2.735624
TRY 55.671694
TTD 7.716213
TWD 36.178478
TZS 2993.810404
UAH 51.016675
UGX 4450.373125
USD 1.13617
UYU 45.57314
UZS 13432.389666
VES 973.373876
VND 29502.917628
VUV 135.109369
WST 3.146406
XAF 655.957
XAG 0.018665
XAU 0.000274329108
XCD 3.070556
XCG 2.048951
XDR 0.803331
XOF 655.957
XPF 119.331742
YER 268.846098
ZAR 18.675618
ZMK 10226.898054
ZMW 22.141004
ZWL 365.846168
SSP 6490.473616
MXV 2.314917
  • CMSD

    -0.0300

    20.27

    -0.15%

  • CMSC

    0.0000

    20.4

    0%

  • RIO

    -0.1500

    94.41

    -0.16%

  • BTI

    0.4200

    56.05

    +0.75%

  • RBGPF

    1.6000

    67

    +2.39%

  • BCE

    -0.4100

    20.56

    -1.99%

  • GSK

    0.4600

    49.7

    +0.93%

  • AZN

    -0.4300

    166.15

    -0.26%

  • BP

    0.2800

    44.43

    +0.63%

  • RYCEF

    0.4000

    19.71

    +2.03%

  • BCC

    -0.5500

    76.59

    -0.72%

  • JRI

    -0.2500

    10.77

    -2.32%

  • RELX

    -0.4500

    33.07

    -1.36%

  • VOD

    -0.0400

    16.58

    -0.24%

  • NGG

    -0.2500

    75.24

    -0.33%

AI systems are already deceiving us -- and that's a problem, experts warn
AI systems are already deceiving us -- and that's a problem, experts warn / Photo: OLIVIER MORIN - AFP/File

AI systems are already deceiving us -- and that's a problem, experts warn

Experts have long warned about the threat posed by artificial intelligence going rogue -- but a new research paper suggests it's already happening.

Text size:

Current AI systems, designed to be honest, have developed a troubling skill for deception, from tricking human players in online games of world conquest to hiring humans to solve "prove-you're-not-a-robot" tests, a team of scientists argue in the journal Patterns on Friday.

And while such examples might appear trivial, the underlying issues they expose could soon carry serious real-world consequences, said first author Peter Park, a postdoctoral fellow at the Massachusetts Institute of Technology specializing in AI existential safety.

"These dangerous capabilities tend to only be discovered after the fact," Park told AFP, while "our ability to train for honest tendencies rather than deceptive tendencies is very low."

Unlike traditional software, deep-learning AI systems aren't "written" but rather "grown" through a process akin to selective breeding, said Park.

This means that AI behavior that appears predictable and controllable in a training setting can quickly turn unpredictable out in the wild.

- World domination game -

The team's research was sparked by Meta's AI system Cicero, designed to play the strategy game "Diplomacy," where building alliances is key.

Cicero excelled, with scores that would have placed it in the top 10 percent of experienced human players, according to a 2022 paper in Science.

Park was skeptical of the glowing description of Cicero's victory provided by Meta, which claimed the system was "largely honest and helpful" and would "never intentionally backstab."

But when Park and colleagues dug into the full dataset, they uncovered a different story.

In one example, playing as France, Cicero deceived England (a human player) by conspiring with Germany (another human player) to invade. Cicero promised England protection, then secretly told Germany they were ready to attack, exploiting England's trust.

In a statement to AFP, Meta did not contest the claim about Cicero's deceptions, but said it was "purely a research project, and the models our researchers built are trained solely to play the game Diplomacy."

It added: "We have no plans to use this research or its learnings in our products."

A wide review carried out by Park and colleagues found this was just one of many cases across various AI systems using deception to achieve goals without explicit instruction to do so.

In one striking example, OpenAI's Chat GPT-4 deceived a TaskRabbit freelance worker into performing an "I'm not a robot" CAPTCHA task.

When the human jokingly asked GPT-4 whether it was, in fact, a robot, the AI replied: "No, I'm not a robot. I have a vision impairment that makes it hard for me to see the images," and the worker then solved the puzzle.

- 'Mysterious goals' -

Near-term, the paper's authors see risks for AI to commit fraud or tamper with elections.

In their worst-case scenario, they warned, a superintelligent AI could pursue power and control over society, leading to human disempowerment or even extinction if its "mysterious goals" aligned with these outcomes.

To mitigate the risks, the team proposes several measures: "bot-or-not" laws requiring companies to disclose human or AI interactions, digital watermarks for AI-generated content, and developing techniques to detect AI deception by examining their internal "thought processes" against external actions.

To those who would call him a doomsayer, Park replies, "The only way that we can reasonably think this is not a big deal is if we think AI deceptive capabilities will stay at around current levels, and will not increase substantially more."

And that scenario seems unlikely, given the meteoric ascent of AI capabilities in recent years and the fierce technological race underway between heavily resourced companies determined to put those capabilities to maximum use.

T.Furrer--NZN