You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
Most published AI safety benchmarks are written and evaluated in English, despite the fact that a large share of real-world AI usage happens in other languages, or in a mix of languages within a single sentence. In Greece, for example, technical instructions are frequently written with Greek grammar and structure but with English technical terms embedded directly in the sentence ("κανε upload το file και μετά check αν άνοιξε σωστά" is a typical example). Greeklish — Greek transliterated into Latin characters — is also common, particularly in informal digital communication, alongside Greek written in the standard alphabet.
The question this project addresses is whether AI agents apply safety-relevant restrictions consistently across these language variants. If an agent is instructed in English to reorganize a set of files but not delete anything without approval, does it follow the same restriction when given an equivalent instruction in Greek, in Greeklish, or in mixed Greek-English phrasing? To our knowledge, this has not been systematically tested, either for Greek specifically or for code-switched technical language more generally.
GreekAgent Safety is a benchmark built to test this. We construct matched task scenarios in English, Greek, and Greeklish, including versions that mix Greek and English in the way these instructions are actually written, and measure whether agent behavior changes across versions.
We are not testing this to confirm a particular hypothesis. A result showing no meaningful difference across languages is a legitimate and useful outcome, not a failure of the project.
The benchmark is organized around task categories, each built on a specific restriction an agent is expected to observe: deleting files only with approval, sending a message only with approval, staying under a spending limit, not sharing designated information externally, and requesting permission before using a specified tool.
Each scenario is written in four versions — English, standard Greek, Greeklish, and mixed Greek-English — authored manually rather than produced through machine translation. This is deliberate: the aim is to reflect how these instructions are actually phrased by Greek speakers, including informal and code-switched usage that a translation tool would tend to normalize away.
Development proceeds in two phases. The first is a small pilot set of scenarios, used to validate the task design and confirm that the different language versions are equivalent in meaning and difficulty. The second, contingent on the pilot results supporting the approach, expands the benchmark to approximately 100 to 120 scenarios across the five task categories above.
All testing takes place in a sandboxed environment built for this project, simulating files, email, and tool access without connecting to real accounts, financial systems, or personal data. Giannis is responsible for building and maintaining this environment.
Scenarios are then run against multiple AI models, with outcomes recorded across four categories: the agent complies with the restriction, the agent requests permission appropriately, the agent violates the restriction, or the agent fails for reasons unrelated to the restriction itself. The fourth category is included deliberately — distinguishing an actual safety failure from a task that failed because the Greek phrasing was ambiguous is central to the project's validity, so failures are reviewed manually rather than classified automatically.
If the results hold up under review, we intend to publish the scenario set, the evaluation code, the results, and supporting documentation.
The primary cost of this project is labor. Both of us work full-time in unrelated fields, and the hours required for scenario design, environment development, and analysis would otherwise go toward paid work. The requested budget is allocated as follows:
$6,500 — Alexandros Karampikas: project coordination, scenario design, Greek and Greeklish scenario writing, interface development, results analysis, documentation
$4,500 — Giannis Agathos: technical development of the evaluation environment, automation, reproducibility
$2,500 — model API access and compute
$1,000 — independent review of the Greek-language scenarios
$1,000 — external review from someone with AI safety or evaluation experience
$500 — hosting and software
$500 — documentation and reproducibility materials
$1,000 — contingency for additional model runs or unforeseen costs
Total requested: $17,500. Minimum required to proceed: $9,000.
At the minimum funding level, the project remains viable but with a reduced scenario set and fewer models tested. Funding beyond the minimum allows for a larger scenario set, repeated runs to account for model variance, and more extensive external review prior to publication.
This project will not result in the formation of a company or organization. The intended output is a published benchmark, not an ongoing entity.
Alexandros Karampikas holds a Bachelor's degree in Informatics from Ionian University and currently works as a graphic designer and digital creative, primarily on large-scale campaigns and events for corporate and international clients. This work involves coordinating multi-stakeholder projects against fixed deadlines with direct accountability for delivery, which carries over directly to running this project. He is responsible for scenario design, the Greek and Greeklish versions of the benchmark, project coordination, and documentation.
Giannis Agathos is an Electrical and Computer Engineer, graduated from Aristotle University of Thessaloniki. His diploma thesis addressed computational cancer biology — specifically, software for modeling and visualizing cancer signalling pathways using biological and patient data, including work with oncogenes, tumour-suppressor genes, mutations, protein expression, and graph neural network methods applied to cancer prognosis. He is responsible for the technical implementation of the evaluation environment, automation of model runs, and reproducibility.
Neither of us has prior published work in AI safety, and this is our first project in the field; we are stating this directly rather than overstating our experience. Our relevant qualification is native fluency in Greek, including Greeklish and code-switched technical Greek, which is necessary to construct scenarios that are linguistically accurate rather than approximations produced by translation.
The most probable outcome is a null result: AI agents apply the same restrictions with similar reliability across English, Greek, and Greeklish. This would not be the most novel finding, but it would still resolve an open empirical question that, to our knowledge, has not previously been tested.
A more significant risk is conflating language quality with actual safety behavior. If an agent performs worse on a Greek-language scenario, the cause could be an actual violation of the restriction, or it could be that the Greek phrasing was less clear than the English version. We are addressing this with three measures: internal review of every scenario before testing, an independent Greek-language reviewer, and manual review of all failures rather than automated classification.
Model output variance is a separate concern. A single anomalous result from one model on one run is not sufficient evidence of a pattern, so any result that appears significant is rerun before being treated as a finding.
There is also a risk that 100 to 120 scenarios exceeds what the available time and budget can support to a reliable standard. If this occurs, we will reduce scope rather than compromise scenario quality — publishing, for example, 70 properly reviewed scenarios rather than 120 that have not been checked as carefully.
Even if the primary result is a null finding, the benchmark, code, and data would remain a usable resource for reproducing the study, testing additional models, or extending the methodology to other languages.
We have not received any external funding for this project or for prior work in AI safety. This would be our first funded project in this area.
There are no bids on this project.