Zürcher Nachrichten - AI systems are already deceiving us -- and that's a problem, experts warn

EUR -
AED 4.248105
AFN 75.770596
ALL 92.830359
AMD 423.156353
ANG 2.0703
AOA 1061.880869
ARS 1706.354669
AUD 1.634265
AWG 2.083564
AZN 1.971037
BAM 1.956743
BBD 2.328663
BDT 141.919628
BGN 1.962242
BHD 0.436037
BIF 3456.475084
BMD 1.156733
BND 1.478902
BOB 13.474494
BRL 6.040925
BSD 1.156227
BTN 110.266145
BWP 15.57645
BYN 3.516708
BYR 22671.957183
BZD 2.325221
CAD 1.60514
CDF 2629.253412
CHF 0.941015
CLF 0.026895
CLP 1058.51481
CNY 7.800083
CNH 7.801606
COP 3644.771598
CRC 520.182191
CUC 1.156733
CUP 30.653411
CVE 110.322941
CZK 24.210763
DJF 205.574957
DKK 7.473191
DOP 67.676714
DZD 152.325697
EGP 57.668157
ERN 17.350988
ETB 187.022939
FJD 2.56037
FKP 0.857112
GBP 0.855002
GEL 3.019524
GGP 0.857112
GHS 12.660102
GIP 0.857112
GMD 85.024294
GNF 10156.334049
GTQ 8.821328
GYD 241.836558
HKD 9.077054
HNL 30.994814
HRK 7.534035
HTG 151.232041
HUF 363.017812
IDR 20622.34285
ILS 3.418265
IMP 0.857112
INR 110.600806
IQD 1514.621659
IRR 1590030.052589
ISK 142.166837
JEP 0.857112
JMD 183.099329
JOD 0.820169
JPY 184.272161
KES 149.519689
KGS 101.156702
KHR 4678.254774
KMF 493.925187
KPW 1041.059598
KRW 1638.477341
KWD 0.357061
KYD 0.963523
KZT 536.501043
LAK 26093.576298
LBP 103531.599032
LKR 384.738715
LRD 209.853842
LSL 18.703876
LTL 3.415531
LVL 0.699696
LYD 7.36193
MAD 10.723725
MDL 20.048876
MGA 4977.398953
MKD 61.563397
MMK 2428.460304
MNT 4161.151475
MOP 9.344704
MRU 46.43078
MUR 54.486429
MVR 17.871955
MWK 2004.766235
MXN 19.69164
MYR 4.726298
MZN 73.927211
NAD 18.703715
NGN 1572.659254
NIO 42.552716
NOK 10.922567
NPR 176.436714
NZD 1.963226
OMR 0.441083
PAB 1.156157
PEN 3.899682
PGK 5.193773
PHP 71.098608
PKR 321.135211
PLN 4.306226
PYG 6939.34455
QAR 4.21485
RON 5.236186
RSD 117.376006
RUB 97.377191
RWF 1700.222344
SAR 4.346103
SBD 9.310058
SCR 15.912008
SDG 694.622124
SEK 11.020235
SGD 1.479697
SHP 0.856984
SLE 28.344188
SLL 24256.10146
SOS 660.758451
SRD 43.926343
STD 23942.02751
STN 24.513298
SVC 10.116313
SYP 15039.835978
SZL 18.702146
THB 38.337629
TJS 10.676846
TMT 4.060131
TND 3.389734
TOP 2.785134
TRY 55.381697
TTD 7.833249
TWD 37.04043
TZS 3065.33888
UAH 51.722311
UGX 4295.33015
USD 1.156733
UYU 46.32453
UZS 13763.752692
VES 890.810236
VND 30246.82002
VUV 137.202351
WST 3.162108
XAF 656.273303
XAG 0.017873
XAU 0.000264
XCD 3.126128
XCG 2.08373
XDR 0.81787
XOF 656.307363
XPF 119.331742
YER 274.381102
ZAR 18.703491
ZMK 10411.984809
ZMW 21.850531
ZWL 372.467396
  • RBGPF

    -0.8200

    71.34

    -1.15%

  • CMSC

    -0.0250

    21.45

    -0.12%

  • CMSD

    -0.0100

    21.58

    -0.05%

  • NGG

    -0.1500

    81.05

    -0.19%

  • RIO

    -0.4100

    95.68

    -0.43%

  • RELX

    -0.2400

    34.43

    -0.7%

  • BCE

    0.1500

    23.47

    +0.64%

  • RYCEF

    0.0800

    20.79

    +0.38%

  • BCC

    -0.8900

    83.24

    -1.07%

  • GSK

    -0.4785

    49.52

    -0.97%

  • BTI

    -0.2900

    57.06

    -0.51%

  • JRI

    0.0635

    12.61

    +0.5%

  • AZN

    -0.7900

    156.45

    -0.5%

  • BP

    0.2196

    42.53

    +0.52%

  • VOD

    0.2000

    16.42

    +1.22%

AI systems are already deceiving us -- and that's a problem, experts warn
AI systems are already deceiving us -- and that's a problem, experts warn / Photo: OLIVIER MORIN - AFP/File

AI systems are already deceiving us -- and that's a problem, experts warn

Experts have long warned about the threat posed by artificial intelligence going rogue -- but a new research paper suggests it's already happening.

Text size:

Current AI systems, designed to be honest, have developed a troubling skill for deception, from tricking human players in online games of world conquest to hiring humans to solve "prove-you're-not-a-robot" tests, a team of scientists argue in the journal Patterns on Friday.

And while such examples might appear trivial, the underlying issues they expose could soon carry serious real-world consequences, said first author Peter Park, a postdoctoral fellow at the Massachusetts Institute of Technology specializing in AI existential safety.

"These dangerous capabilities tend to only be discovered after the fact," Park told AFP, while "our ability to train for honest tendencies rather than deceptive tendencies is very low."

Unlike traditional software, deep-learning AI systems aren't "written" but rather "grown" through a process akin to selective breeding, said Park.

This means that AI behavior that appears predictable and controllable in a training setting can quickly turn unpredictable out in the wild.

- World domination game -

The team's research was sparked by Meta's AI system Cicero, designed to play the strategy game "Diplomacy," where building alliances is key.

Cicero excelled, with scores that would have placed it in the top 10 percent of experienced human players, according to a 2022 paper in Science.

Park was skeptical of the glowing description of Cicero's victory provided by Meta, which claimed the system was "largely honest and helpful" and would "never intentionally backstab."

But when Park and colleagues dug into the full dataset, they uncovered a different story.

In one example, playing as France, Cicero deceived England (a human player) by conspiring with Germany (another human player) to invade. Cicero promised England protection, then secretly told Germany they were ready to attack, exploiting England's trust.

In a statement to AFP, Meta did not contest the claim about Cicero's deceptions, but said it was "purely a research project, and the models our researchers built are trained solely to play the game Diplomacy."

It added: "We have no plans to use this research or its learnings in our products."

A wide review carried out by Park and colleagues found this was just one of many cases across various AI systems using deception to achieve goals without explicit instruction to do so.

In one striking example, OpenAI's Chat GPT-4 deceived a TaskRabbit freelance worker into performing an "I'm not a robot" CAPTCHA task.

When the human jokingly asked GPT-4 whether it was, in fact, a robot, the AI replied: "No, I'm not a robot. I have a vision impairment that makes it hard for me to see the images," and the worker then solved the puzzle.

- 'Mysterious goals' -

Near-term, the paper's authors see risks for AI to commit fraud or tamper with elections.

In their worst-case scenario, they warned, a superintelligent AI could pursue power and control over society, leading to human disempowerment or even extinction if its "mysterious goals" aligned with these outcomes.

To mitigate the risks, the team proposes several measures: "bot-or-not" laws requiring companies to disclose human or AI interactions, digital watermarks for AI-generated content, and developing techniques to detect AI deception by examining their internal "thought processes" against external actions.

To those who would call him a doomsayer, Park replies, "The only way that we can reasonably think this is not a big deal is if we think AI deceptive capabilities will stay at around current levels, and will not increase substantially more."

And that scenario seems unlikely, given the meteoric ascent of AI capabilities in recent years and the fierce technological race underway between heavily resourced companies determined to put those capabilities to maximum use.

T.Furrer--NZN