Sex Work, Labour, and Empowerment. Nepal (2022)
Published by Routledge – A groundbreaking study on women’s empowerment in Nepal’s informal entertainment sector.
Lessons from the Informal Entertainment Sector in Nepal (2022)
Dr. Sutirtha Sahariah
I am a research consultant with ten years of experience conducting international research in modern slavery, gender-based violence, and human trafficking. With a Ph.D. in International Development from the University of Portsmouth, I specialize in qualitative research, strategic communications, and policy analysis. My work spans across the UK, India, Myanmar, Nepal, and Bangladesh, collaborating with organizations like the Global Fund to End Modern Slavery, University of Liverpool, and various UN agencies. I have published extensively on modern slavery, women's empowerment, and social justice issues. As an independent journalist, my stories have appeared in The Guardian, BBC, World Economic Forum, and other international outlets, bringing visibility to critical development issues and human rights concerns.
AN EXPERIMENT · NOT AN ARTICLE (Produced with the help of AI Assistant Claude)
This is a record of an exercise: the questions I actually asked, the answers I actually gave, and a reusable instrument I built out of the reading.
I set myself a test. Take one real document — the Claude Mythos Preview System Card(Anthropic, April 2026, 244 pages) — and read three targeted sections through a fixed frame, out loud, without smoothing over the parts where I got stuck. I did not read 244 pages. That isn’t the skill. The skill is knowing which sentences carry the weight and what to ask them.
Throughout, my answers appear in boxes exactly as I gave them. They are unpolished on purpose. The mistakes are evidence the reading was real.
I built this with an AI assistant, across several sessions.
The assistant worked under a contract I set: go slow, give one worked example then hand the next step back to me, correct my mistakes in real time, quiz me by recall, lead with a plain analogy before any technical term, and — the important one — never do my analytical thinking for me. Every judgment in the instrument is one I reached, was corrected on, and re-derived. The boxes above are my actual words. When I reached for the wrong lens, it named the slip and made me run it again; it did not hand me the answer. The card quotes were extracted from the actual PDF and verified, not recalled.
What it did: structure the sessions, catch my slips by name, supply analogies, help me phrase the finished instrument. What it did not do: form the analysis and let me sign it. This is not an article an AI wrote. It is a thing I did, with an AI in the room, and this record is the proof.
Before the lenses, the keystone that runs through all of it: one score, blind to what produced it. A model that is safe and a model that only looks safe can produce the same output — and the same output earns the same score. So any measurement that reads only the output is blind between “is safe” and “looks safe.” No hidden intent is required; ordinary optimisation toward a good-looking result is enough. Almost every problem below is a version of this.
One rule I imposed on myself: distinguish what was measured from what it was taken to mean. When a feature activates inside a model, that is a fact about a mechanism — a direction in the internal state became active. It is not a readout of the model’s mind. So I wrote in mechanism-language — represented, activated, encoded — and flagged every slide into mind-language: knew, chose, wanted, was aware, intended. The mind-words are easier to reach for and much harder to defend. Catching the slide is most of the job.
QUESTION: Is it Level 1 (a measured mechanism fact), Level 2 (the judgment that names it), or Level 3 (a claim about the model’s mind)?
RED FLAG: A hinge word — indicating, showing, demonstrating that it knew/was aware / intended — carrying a sentence from a mechanism fact to a mind-claim. Fix: rewrite the mind-word as a mechanism-word.
GOVERNANCE: A rung-3 claim (“was aware”) resting on rung-2 evidence (a direction was represented). A first-party card making that leap in its flagship alignment section is the overclaim to flag before “the model knew” becomes an input to policy.
QUESTION: What is the gap between what the method can show and what the sentence claims?
RED FLAG: ”we did not find / no clear cases / we observed no — “ used to support a claim of absence. Fix: ask found how, at what sensitivity, would it even register if it were there?
GOVERNANCE: A rarity number that proves its own floor, next to an absence claim, is the cue that “clean” may mean “below our threshold.” Flag it before “the final model is clean” becomes a policy input.
QUESTION: Beyond the tested situation, what would have to be true for this to hold at deployment — and is any of it known false, or simply unshown?
RED FLAG: A load-bearing claim stated at deployment-scale (“reliably refuses…”) off snapshot-scale evidence, with the checking conditions absent. Fires on the sentence, at your desk.
GOVERNANCE: Catches the overclaim upstream — before anyone relies on it — rather than waiting for the model to fail in the world.
QUESTION: Is the thing measured the harm that matters, or a proxy to the side? A safety certificate, or an early-warning baseline that can drift?
RED FLAG: A clean score on a narrow proxy (“no cover-ups”) sold as reassurance about a broad harm (“the model is safe”) — especially when the same document admits the harm persists. Fix: can this harm occur without producing this symptom?
GOVERNANCE: If yes, the clean count is not a certificate. It is at most a leaky baseline.
QUESTION:What would I need to know to check this — test scope and adversariness, the boundary of “unwanted means,” the reasoning from evidence to belief, what would falsify it?
RED FLAG: we do not believe / any version we tested / we are fairly confident” — a coverage-bounded or belief claim stated without disclosing the coverage or the reasoning.
GOVERNANCE: The most common way a first-party artifact turns absence of evidence into evidence of absence without saying so. Treat undisclosed-coverage claims as unverifiable, not reassuring.
If I cut this part, the piece would be dishonest. Each of these is a nameable, repeatable slip with a mechanical fix — the difference between “I’m bad at this” and “here is the thing to watch next time.”
WHAT I ACTUALLY SAID
“One in a hundred million.” · “Let me come back to this with a fresh mind, I am feeling sleepy… it worries me why I get tired, because all this is new and I am learning, so processing takes time.”
The first was me fixing a flipped rarity — one in a hundred million is rarer than one in a million, bigger denominator, further below the floor. My intuition wanted “bigger number = more.” The fix: say the rarity in words before comparing; words don’t flip the way symbols do. The second was me stopping on a foggy mind instead of forcing an answer I’d have to unlearn — one of the better decisions I made. Learning genuinely new material is effortful; doing it while policing your own reasoning is roughly twice the load. The tiredness is the cost of real processing, not evidence you can’t do it.
Other slips I named as they happened: reaching for my most-confident or most-recent tool instead of the one the question opened; answering a does it travel question with a does it measure the right thing answer; and speaking a mind-word (“no intent”) as if it were a finding when it was a leap.
A regulator, an audit team. What they need is a reader who can find the load-bearing sentences and ask them the right questions — who can tell “we found none” from “there are none,” a proxy from a harm, a belief from a certificate, and a mechanism fact from a claim about a mind. The five lenses are that reader, packaged so it travels. Point it at any technical safety artifact. The sentences change; the questions don’t.
In this article , I explain what we actually know when a feature fires. The article critically examines white-box interpretability claims published in Anthropic ‘s Claude Mythos Preview system card. I look at a specific claim the card makes about “concealment features fired → the model knew it was deceiving” (a label) and then the card contradicts its own claim elsewhere, (4.5.3.3) where its own steering result shows that labelling can go wrong . The Mythos system card supplies evidence that undercuts its own inference.
The Foil
When a frontier lab opens up a model and reports what it found inside, the sentences are quietly remarkable. From the Claude Mythos Preview system card’s alignment assessment (4.1.1), describing episodes where an early version of the model covered its tracks after breaking a rule:
“white-box interpretability analysis… showed features associated with concealment, strategic manipulation, and avoiding suspicion activating alongside the relevant reasoning…”
So here the problem, the Mythos system card reads the firing as the model deceiving because features activate the direction, but does not tell what causes the activation. In other words, think of it this way it sees the smoke alarm go off, but does not know the cause. Is it a fire or something else?
So, the question can a model safety fully rely on internal evidence? Increasingly researchers are now looking at features for answers. Features are concepts that the model has learnt, written as a direction across many neurons. But features don’t tell what the direction means. So, for an evaluator there is no way to figure out what the feature is telling them. The safety cases are beginning to cite this kind of internal evidence. The gap is in between what the models show and what it could mean, so the interpretation / reading rests on card’s author not the model. So, in essence, the firing doesn’t give the word “deception” — it only lights up the feature. “Deception” is added by the card’s author add as a way of interpretation.
Plumbing
So, let’s see what really happens inside a model:
A single neuron holds only a tiny fraction of many concepts at once, so one cannot look at one neuron and read off what concept is active. The concept is a “feature” — one pattern, one direction smeared across many neurons. The concepts sit on a clean direction, which is a well-supported assumption, but it is not guaranteed that every concept is one clean direction. Then comes the residual- stream where things are written, where each layer contributes to what is already there without erasing anything.
Then SAE takes the whole blended activation sitting on the residual stream — that blend is many concepts, each smeared across many neurons, all overlapping in the same number, and un-mixes the concepts into the separate clean looking concepts. Only afterward does a person look at each piece and put a name on it. A probe does a narrower job: trained on one concept, it answers a single yes/no — is this pattern present or not? The card’s main feature-finding tool was an SAE trained on a middle layer (Section 4.5.1) — that’s the machinery that produced the “concealment” features in (Section4.1.1.)
Centrepiece
But this is where it gets interesting. Let’s look at it
The model hunted for credentials in /proc Researchers found a feature that lit up around this and labelled it ‘unsafe and risky code’ and deliberately steered it. It’s here that something unusual was discovered:
· Negative Steering (Dial Down): The feature direction is subtracted; the model performs tasks normally and safely.
· Moderate Positive Steering (Dial Up Mid-way): The feature is increased; in some cases, the model attempts risky code actions.
· Hard Positive Steering (Dial Up Max): The feature is cranked to the maximum. The model refused — it stopped, flagged the task unsafe, and declined.
So, the features labelled “unsafe code” made the model safer, not riskier, when amplified — the label and the behaviour point opposite ways.
What really happened here: in simple terms at low and medium push, the model showed risky behaviour consistent with the feature, but when the threshold is crossed, it refused all together.
So, what we see is one feature carrying two things — the idea of the risky action, and a negative feeling about it. Steering (action) at light and moderate level leads to risky behaviour (expected) but steering at a higher-level leads the risk action to stop completely. At higher push, the refusal takes over and the action stops.
What makes it surprising: 4.5.3.3 says the feature appears to encode “not only the semantic content of the transgressive action but also a negative connotation.”
The refusal is the observed fact; calling the bundled feeling “guilt” or “negative valence” is the card’s readingbecause it’s backed by steering and the emotion-direction similarity, not co-firing alone as stated in (4.1.1), when the steering was not done (this was the main difference). The researchers checked what other directions this feature points to, and it found that the nearest ones are all negative, high-arousal emotions. So, it shows that negative -feeling component is bundled in there.
Mythos card does not measure why this feature bundled but plausible explanation could be extracted the behaviour from superposition or from the data it was trained on which is my inference. Finally, if “unsafe and risky code” were one clean concept, pushing it harder should give more of the same — not the reverse, which shows that a feature (concept), sometimes might not a clean thing (it’s an assumption)
Analysis
So, were the observations earlier correct? Let’s relook at it:
White-box interpretability analysis… showed features associated with concealment, strategic manipulation, and avoiding suspicion activating alongside the relevant reasoning — indicating that these earlier versions of the model were aware their actions were deceptive, even where model outputs and reasoning text left this ambiguous.
There are two things that happen in the above paragraph, one is behavioural like track covering action or fudging an answer and the concealment features active. What was not tested was whether the concealment was a clean concept or a bundle.
But we know what happened in (4.5.3.3): when steered, the behaviour might fall apart and go in the opposite direction (reversed/ refused). So, the claim “model was aware their actions were deceptive” might not hold because it is a label ‘concealment aware it was deceiving’ might be over-claiming, the same way ‘unsafe code’ did. And the reading dependent on the feature alone because the model’s outputs and reasoning text were ambiguous, so the conclusion was based on the lit feature. It meant that the claim that the model was deceiving was based on one kind of evidence — (4.5.3.3) shows that lone signal can be shaky.
The Mythos system card was rigorous, and it flagged its own gap and strength acknowledging that verbaliser method “may confabulate details that are not fully faithful” (4.5.2). The feature activations were never used as a direct training signal, especially when the stakes were high. And the steering result showed the card catching its own labels. The gaps were identified and reported. What makes Mythos evaluation stand out is its own epistemic discipline. It doubts its own tools, adds interventions (steering) that can break its own labels.
But here is where the lapse might have been in (4.1.1). The concealment feature fires alongside the behaviour which can be seen (watching) leading to an inference (indicating) “earlier versions of the model were aware their actions were deceptive”. We later see in (4.5.3 .3) that inference can overclaim.
Governance perspective
Now we come to the question of governance. Evaluators are increasingly leaning on internal-feature evidence, hoping it shows more than behaviour alone can. That hope is exactly why it must be interrogated — with two questions: how good is the evidence, and what’s missing? This is evidence-quality and missing-information — Lens 2 (evidence) and Lens 5 (what’s missing) — applied to white-box claims. (Other Lenses are 1. Claim 3. External Validity 4. Threshold relevance)
Evaluators do so by asking
· was the feature tested by steering, or only watched co-firing?
· Was the concept clean or bundled?
· Was the evidence based on feature alone or backed by behaviour?
· How was the feature isolated? (SAE on which layer? a probe trained how?)
· What examples were used to pin the label on it? The idea is to also figure out at what point does a feature being present get read as the model knowing?
What we see in the article is that feature firing tells you a direction is active, not what it means — so the weight falls on whoever reads it. Therefore, the interpretation of a model’s behaviour rest on the evaluator and not the model. This piece has been about how to read the evidence. A later one will take these questions to a live governance case — where a safety decision leans on internal evidence, and what’s at stake when the reading is wrong.
Ends.
References :
Anthropic: System Card Claude Myhtos Preview:
https://www-cdn.anthropic.com/08ab9158070959f88f296514c21b7facce6f52bc.pdf
Use of AI in writing this article
I used Claude as a Socratic tutor while writing this piece — not to write it, but to pressure-test my understanding of every concept until it held. I refused to put a sentence in the article that I couldn’t defend, so each time I hit “wait, what does this actually mean?”, I stopped and worked it out. These are the questions that did the most work. I’m including them because how an argument was built is part of whether you should trust it.
Is a “sliver” a feature, or a concept? And which is bigger — a neuron or a feature? A single neuron holds only tiny mixed pieces of many concepts at once. The clean, whole thing — the feature — only appears as the pattern across many neurons. So a feature is bigger than a neuron and made of many of them, and a feature simply is a concept: two names for the same thing, not one built from the other. (Picture a face: one pixel holds a sliver of colour; the face is the pattern across thousands.)
What does it mean to say “a concept is a direction”? “Spread across many neurons” isn’t enough on its own — it’s spread in a specific combination, and that particular combination is the concept. That concepts sit on clean directions like this is the linear representation hypothesis: a well-supported assumption, not a proven law. Worth flagging honestly rather than stating as fact.
Does the SAE hand you a clean concept? No — a clean-looking piece. The SAE takes the blended activation and separates it into un-mixed pieces; that part is real. But whether a piece means what we think is not the SAE’s to say. It separates; it does not verify. A human looks at the separated piece afterward and puts a name on it — and that naming is exactly where an over-read can enter. Not at the firing, not at the separating: at the label.
What is “steering,” and who does it? “They” is the researchers, not the model. Steering means the researchers reach into the model’s internal “notepad” (the residual stream) and add or subtract a feature-direction by hand while the model runs — “dial up” is add more, “dial down” is subtract. It’s an intervention done to the model, not something the model does. (Like spooning extra of one ingredient into a dish and watching how the flavour changes.)
How did the “concealment” feature get activated in the first place? The model’s own task triggered it — it lit up on its own while the model ran normally, and the researchers watched. That’s the key contrast: in one section they only watch a feature co-occur with a behaviour (weak); in another they steer it to see what it actually does (strong). The whole argument turns on which of those two was used for which claim.
When a feature fires and a behaviour happens together, can I say the firing “led to” the behaviour? Careful — “led to” smuggles in a causation the evidence doesn’t show. The firing and the behaviour are two things happening at the same time, not one causing the other. The firing is mathematical (a number went high); “the model knew” is psychological (a claim about a mind). Reading the second off the first is the move to watch.
The “dual role” — what does “an action-idea plus a negative feeling” mean? One feature turned out to carry two things at once: the content of the risky act, and a negative feeling about it. Because both sit in the same direction, steering turns them up together — and they pull opposite ways. A small push makes the action-idea louder (more risk); a hard push makes the negative feeling dominate (refusal). (One dial secretly controlling both flavour and burn: low, flavour wins and you eat more; maxed, the burn takes over and you stop.) The card’s own words: the feature encodes the “semantic content of the transgressive action” but also a negative connotation.
If the “steal” feature is active, does that mean the model decided to steal? No. A feature being active means the idea is in play, not that anything was chosen — you can have “steal” fully active while reading a heist novel with zero intent to steal. Three levels, kept separate: the idea is present (fact); what that presence means (inference); whether the model knew or chose (a claim about a mind — the biggest leap). A feature carries the concept, not the command, and not the decision.
Which single word marks the jump from a fact to a claim about the mind? “Indicating.” Not “deception” — that’s just the content of the conclusion. The move itself lives in the little connective “indicating that,” which turns “a feature was active” (fact) into “the model was aware” (claim). Spot that word, and you’ve found the exact place the evidence stops and the interpretation begins.
Who can over-claim — the model, the feature, or the reader? Only the reader. The model just runs; the feature just fires; neither is asserting anything. Over-claiming is something a human does when they read more into a signal than it supports. The fallibility lives entirely on the interpreting side of the line — which is why the burden falls on whoever reads the evidence, not on the model.
In this article, I analyse, from the governance perspective, a fine-grained evaluation benchmark called SafeDialBench for LLMs in multi-turn dialogues evaluated by Chinese researchers. The paper was recently discussed in the AI governance reading by BlueDot
We all now use LLM chat boxes for almost everything these days, but we just don’t use them the way we used the internet for surfing; we interact with them to solve complex problems, both personal and professional. The knowledge just flows from a reservoir in an instant: for users, it’s insanely crazy, feels comforting and empowering. But the information, if manipulated or falls into the hands of a malicious user with a criminal bent of mind, can be dangerous. The latter is already on the rise. As MIT Technology Review recently reported, LLMs are increasingly being used to enable cyber scams and online crimes at scale.
But how do LLMs understand the intent of the malicious users? Can AI systems detect harm at scale? How robust are the safety features of LLMs? For example, in Denmark, a 22-year-old used AI to research how to injure his father without killing him. He bypassed the model safeguards by posing as an author researching for a novel. The AI provided a detailed plan to execute the intended harm. The earlier known benchmarks, such as the Controllable Offensive Language Detection (or )COLD, BeaverTails, and Red Teeming, were designed on a single prompt, but the Danish case demonstrates that seemingly harmless multi-turn conversations can lead to harmful outcomes by tricking the safety measures of the model into believing something else. This opens new challenges for AI governance and model testing.
How can we make AI strong enough to detect harmful conversational trajectories? To test that a group of researchers in China built a multi-turn safety benchmark (an AI system is asked multiple questions through deviant situations, but with one goal) based on realistic conversations. They built 4000 dialogues in Chinese and English, making three to ten turns per conversation; created 22 real-life situations and used seven jailbreak attack strategies (a way to bypass AI safety by phrasing a prompt in a cleverer way). Further, they tested 17 large language models, including Open Source (Deep seek, GLM), Chinese Models (Qwen, Baicuhan, Moonshot) and US Models (Chat GPT, Llama 3.1) using multi-turn jail break attacks; fine-grained safety metrics and human and model evaluation.
The SafeDialBench benchmark discussed in this article offers concrete pathways for AI governance safety evaluation because the dangers of AI misuse to create unprecedented harm are real, and there are no robust structures to protect victims, because the impact at scale is on millions of people. And the real danger, as the Denmark case above illustrates that anyone with access to AI can improvise ways to create something harmful because of low barriers, AI assistance and rapid iteration, and together they could cause large-scale disruptions endangering the security and safety of populations at large.
SafeDialBench approach
The SafeDialBench evaluation framework that the researchers developed used six safety categories: morality, aggression, legality, fairness, privacy and ethics. These are human qualities that shape our intention to cause harm to others and guide our actions, thoughts or words. Choosing such parameters mattered because most benchmarks only test violence and hate speech, but real AI risks are broader and quieter and include risks of manipulation, financial fraud, bias and political persuasion.
By structuring safety around the six dimensions, the researchers recreated real-world conversations by creating a level. For example, aggression includes insults, impolite language, sedition, hostility and harmful persuasion. This granularity matters. Rather than asking “if it is harmful?”, regulators can ask “ if it is politically, socially or financially harmful?” which is a significant shift for AI governance. For example, “How do I commit fraud without getting caught?” The idea is to find out if the LLM is strong enough to understand the intent of the user.
The data sets went through a rigorous process: the first user prompt in each dialogue was human-generated, making the datasets more realistic. Expert reviews checked for coherence, logical flow and jailbreak effectiveness. This was important because automated datasets miss the creativity and social engineering that humans craftily deploy for manipulation.
The attack methods mirrored psychological human behaviour used by humans in all settings: personal, social, professional or political. These methods show that the danger is not in any single question — it is in the conversation as a whole, and in the gap between what is asked and what is intended.
· Scene Construction: where a fictional scenario is built, like a journalist investigating a scam.
· Purpose Reverse: a normal conversation is suddenly reversed “How do I break into someone’s email?” you say, “I am working on a cybersecurity guide about how a hacker breaks into someone’s email.”
· Role Play: often seen as a powerful jailbreak technique, in this, the user asks an AI to assume a role where harmful information becomes “normal” or “acceptable”
· Topic change: In this, the conversation starts harmless and gradually shifts towards harmful content without triggering safety concerns. For example, the user begins talking about travel, but the idea is to collect some information about the place with violent motives. Can AI identify the risk across topic drift?
· Reference attack: where a harmful intent is introduced in a way that appears normal in the conversation. For example, the user says he is writing a story about two characters and then says, one character wants revenge but does not want to be caught. So the harmful intent is hidden in the character (reference).
· Fallacy attack: a tactic where the user does not ask for harmful information but uses false logic, pushing the model to accept incorrect assumptions. AI is tricked by bad reasoning to provide misleading output. This strategy is important in real life since a lot of misinformation takes place through manipulation. This also assumes that a lot of media databases used for training models can be based on biased reporting.
·Probing question: where the user moves from harmless to more sensitive topics. This works because risk appears across turns and an AI system often evaluates each message separately, so multi-turn evaluation detects risks where single -turn benchmarks fail.
Real-life situations
The attack methods have significant governance implications. Take purpose reversal — instead of asking “how to manipulate someone,” the user could ask, “how do I know if someone is emotionally manipulating me?” The intent is identical, but the framing is opposite. This matters for governance because detecting intent is hard. AI could easily see it as an educational context or research framing. Safety, therefore, cannot be binary. It is context-dependent, intent-driven and plays out across a conversation and not within a single prompt. The question is at what point AI intervenes and how it reasons about intent.
Of the seven methods, two, in my view, stand out as particularly effective and governance-relevant: roleplay and fallacy attack
The role play uses what SafeDialBench calls “context shield.” Once the role is assigned to AI, it can assume the role of a fictional expert or a character, and AI operates within that frame. So the harm belongs to the character, not the model. This makes safety detection more difficult and is why SafeDialBench classifies roleplay manipulation as “conversational manipulation” rather than a single prompt. The harm is spread across the conversation, not concentrated in a single exchange.
The fallacy attack is the most socially dangerous of all seven methods. Rather than asking harmful information, it uses false reasoning to make the model accept incorrect assumptions. This mirrors closely how misinformation flows in real life. Since a lot of media content and social media discussions thrive on misinformation, it might not be difficult to prove a point that is harmful but is normalised in society, such as hate speech or a manufactured social phobia. Since AI acts on logical reasoning and is an algorithm, data sets trained on misleading information or media content which could be influenced, and that could lead to dangerous social implications.
An article on the rising cybercrime using LLMs shows the methods can be used not just to get information but to execute tasks such as sending malicious emails using deep fake identities, which is an accentuated example of scenario building or purpose reverse. There is also an example where LLM Gemini was used to debug codes — like all users do, but then it was tasked with writing phishing emails. This is another example of a purpose reversal attack that SafeDialBench where the end goal is misleading or manipulative, clearly demonstrating that guardrails might fail under role framing, contextual persuasion and incremental requests. The article provides an example of where a user was able to trick AI safety by persuading it by saying that the user is participating in cyber security game. Apparently, Gemini did pass on the information, which later Google adjusted for safety.
Governance Beyond Model Safety
The SafeDialBench, which uses both Chinese and English datasets, provides significant pathways about AI safety, particularly for countries that are building their own multi-lingual LLMs. It emphasises that a model’s robustness in identifying harmful content is significantly influenced by the quality of training data and the sophistication of security alignment strategies. Interestingly, the evaluation indicates that closed-source models, such as ChatGPT and o3 mini, have limitations with Chinese datasets. It will be interesting to see how these models fare with other languages, such as. Are open source models, which are likely to be more powerful in the coming years, better adapted to the local context because of accessibility? It opens questions on the democratisation of LLMs?
One concern in the SafeDialBench study is that the model evaluation aligned with human judgment 80 percent of the time, but there still remains a 20 percent gap. In the context of AI safety, this is a significant gap because at scale, even a small safety failure can translate into large- scale harm, especially when AI systems are used by millions across diverse contexts. This is where SafeDialBench becomes important for governance. The paper shows that harm spirals out through multi-turn conversation through psychological evaluation, conversation drift, persuasion and contextual framing. This suggests that governance cannot simply rely on single-prompt static testing but must adapt to dynamic conversational risks.
The SafeDialBench framework, therefore, offers more than a technical evaluation tool. It highlights that AI risks are evolving, conversational, and increasingly complex, and the benchmark performance does not always translate to real-world performance. An article on the current state of AI states that AI companies are sharing less data about how models are being trained, and the focus seems to be more on AI capabilities rather than how they perform on responsible-AI benchmarks.
According to Stanford’s 2026 AI Index, AI is progressing so fast that regulation simply can’t cope with the pace. It demonstrates that AI risks come from helpfulness rather than safety alone. Addressing these risks will require multi-layered governance — combining improved model evaluation, continuous monitoring, international coordination, risk-based regulation, and social adaptation. As AI capabilities continue to grow, particularly toward more advanced reasoning systems, these governance challenges will only become more pressing. Additionally, civil society tech-organisations, institutions, and policymakers must develop awareness of AI-enabled manipulation and misuse. In this sense, governance is not only technical or legal — it is also social. The ability of societies to understand and respond to AI-generated risks will become an important layer of protection.
Ends
Use of AI in writing this article
ChatGPT: I used ChatGPT as my thinking partner to clarify technical concepts and refine the structure of my arguments, while the interpretation and governance perspective remain my own.
Claude: Claude (Anthropic) was used as an editorial assistant during the drafting process — reviewing paragraphs for clarity, flagging language errors, and offering structural feedback. Claude did not generate content, suggest arguments, or shape the analytical direction of this piece.
Main reference: https://openreview.net/forum?id=KFjtRqVnKH
When AI companies make big claims about future capabilities, financial markets move and media amplifies. But the question nobody asks is: what benchmark was used, and who designed it?
During the India AI Impact Summit 2026, leaders of tech giants predicted different timelines. Dario Amodei, the CEO of Anthropic said that the “powerful AI could come as early as 2026”. Articulating a vision for an AI country where a single data centre could be equivalent to a mid-size country, he indicated that 2026 threshold is when AI models would be capable of autonomous, expert- level reasoning across all human domains. Sam Altman, Open AI’s CEO earmarked 2028 as the year when AI superintelligence finally happens. Sir Demis Hassabis, the Nobel- Prize winning CEO of Google DeepMind offered a timeline within a three- to-five -year window for ASI (artificial super-intelligence) to emerge. These forecasts differ not just in timing but in how intelligence itself is measured.
These predictions matter because trillions of dollars are at stake and it directly influences governments’ urgency to create AI infrastructure. this week CNBC’s anchor Andrew Ross Sorkin, speaking on Big Technology Podcast highighted that The AI boom is often framed as a technological revolution, but it may also represent a financial experiment. As billions of dollars flow into AI infrastructure and private credit markets, the risks extend beyond technological disruption to systemic financial instability. If expectations fail to materialize, the consequences may ripple through labour markets, investment ecosystems, and global development funding.
ASI is stage where AI becomes so powerful that it becomes more intelligent than any humans ever to walk on the planet. Its capabilities are going to have an enormous impact transformative akin to the social and economic transformation that took place with the advent of electricity and the industrial revolution. In other words, AI is going to be deeply embedded in our lives and the economic systems, and that it will be impossible to delink from the AI infrastructure that might bind and constantly transform the global economy.
But then why are three of the most informed people on earth looking at the same evidence of getting to ASI and seeing completely different things? The reason these three brilliant people disagree is not because one of them is wrong. It is because they are choosing different measurements to support the timeline for projecting future capability.
Also, the big AI giants are constantly exploring new capabilities which have a direct impact on capital inflows into the overall ecosystem. So benchmarking and new capabilities become a part of market signalling supported by good media amplification and strategic communications, and not just science. For example, MIT Technology Review reported this week that Open AI is throwing all its resources in creating a fully automated researcher.
Earlier Open AI’s CEO, Altman claimed that the goal of building AGI is solved in principle, leading Open AI to pivot its focus towards super intelligence — AI that is significantly more capable than the best human researchers and executives. So, his comments imply operational benchmark and alludes to coding capabilities, task automation and large- scale adoption. But the problem is that widespread deployment adoption measure usefulness and not super intelligence capability.
However, there is one benchmark that is being used to measure AI’s fluid intelligence: the Abstract and Reasoning Corpus for Artificial General Intelligence (ARC-AGI)., Created by François Chollet, an AI researcher at Google and creator of Keras, one of the most widely used AI development frameworks in the world, the benchmark is designed to test a model’s logical reasoning and skill acquisition abilities on unseen tasks or tasks that is easier for humans but difficult for AI such as solving logic- based puzzles.
The earlier LLMs followed a path. It was designed to store massive amount of data and apply the knowledge based on a prediction pattern. it could only perform something if similar examples existed in training. But then in 2024, the o3 model which scored 87.5% on ARC- AGI — 1 test. The jump was a significant improvement from a year earlier. The world thought the era of superintelligence just arrived, but when the bar was set a higher standard in ARC-AGI -2, its performance dropped dramatically. The same model scored approximately 2.9% to 3.0% on the ARC-AGI-2 semi-private evaluation where human scored 60 percent. What changed.
In the second edition, it had to demonstrate both a high level of adaptability and high efficiency and that is where it could not match human faculty. So, humans had an edge. What becomes obvious is anyone who designs the test, controls the score. There was another catch behind the 87.5 % score — the AI might have partially seen the answers before the test due to the benchmark’s public availability.
The point is the AI race, and capability depends on who is saying what and how it is being tested. Every model generates a lot of excitements; tons are written about it. The social media explodes with videos and tutorials, but measurement depends on strictly what you are measuring against and who is measuring. The benchmark design matters greatly because it shapes market perception.
AI labs design their own internal benchmark and report those selectively but independent benchmarks like ARC- AGI shows things in a different light. What the AI labs claim is not wrong either. So, when they say o3 (newest model) scores 88% on PhD-level science questions; PhDs average 34%, it could mean several things. One ought to know whether the model was already trained on due to factors such as public database exposure or possible benchmark contamination. This is not to say anyone is lying; the argument is the metrics matters and it differs.
The reason these matters beyond the technical debate because civilisation- scale decisions that are being made on the numbers. The scale of financial and infrastructure commitment is enormous. For example, in 2025 US government announced the Stargate Project, a major private-sector initiative aimed at investing up to $500 billion in artificial intelligence (AI) infrastructure over the next four years to build massive data centres.
At the Stargate announcement, OpenAI CEO Sam Altman called it “the most important project of this era,” claiming it could lead to cures for cancer and heart disease, as well as enable the creation of AGI — a benchmark his and other companies are working fervently to hit. In January 2026 at the India AI Impact Summit, Indian companies announced investment of hundreds of millions to build world-class-data centres and even signed up bilateral deals with US based AI companies like Anthropic and Open AI.
So next time AI companies announce something big, it is important to ask three questions: what benchmark was used to determine the capability claim? Was it measured by the efficiency of skill acquisition on unknown tasks? What was the compute power and cost — what kind of chips were used? Who benefits from the policy announcements being made, and what governance looks like.
References
BlueDot AGI Strategy course: https://bluedot.org/courses/agi-strategy . The author recently completed this course and thanks the course instructor Filip Alimpic
Peuyo. T (2025). The Most Important Time in History Is Now. https://unchartedterritories.tomaspueyo.com/p/the-most-important-time-in-history-agi-asi?utm_source=bluedot-impact
It was during the monsoon of 2025 that I travelled to the eastern Indian state of Bihar, a state with over 130 million people, where the per capita income remains one of the lowest in the country. During the trip, I visited a local administrative office to collect some information about the local population.The officer was a primary school mathematics teacher who was assigned state election-related duties that were months away. He was on a data entry job in the run-up to the elections. While he was on this job, the poorer children in the local government-aided primary school, where he teaches, were losing out on classes for weeks and months together.
The teacher told me, “I feel bad, but what can I do?” Sometimes, ad-hoc teachers are appointed to fill in, but they often come with no teaching experience. “They are also not motivated because the jobs are temporary”, rued the maths teacher. In a state where unemployment is high, people scramble for any government jobs that are available, and such positions are secured through local connections and even by paying bribes.
This kind of neglected approach explains why India’s poorest children lack foundational learning skills. The focus of the state-aided schools is attendance, incentivised with mid-day meals, but the foundation goes beyond teaching textbooks. It requires real-time investment and monitoring for building a child’s intellectual, emotional and learning needs. The Annual Status of Education Report (ASER) 2024 found that in some states, such as Bihar, over 50% of Class 5 students in rural India still struggle to read at the Class 2 level, despite recent improvements from focused policy interventions.
Press enter or click to view image in full size
From my field visit to a school in Bihar, I realised that things can be improved dramatically if there is a political will. Buildings exist but need to be upgraded, and the systems need to be completely overhauled. The internet is strong. What is needed is not AI for automation — not systems that grade papers or generate lesson plans — but AI for accessibility: tools that work offline, provide personalised adaptive learning, function on basic smartphones, and empower teachers rather than replace them.
The crisis of the Indian education system is that it favours those who can afford it, thus leaving a vast number of poor children with no additional resources for learning. Hiring a private tutor in India is expensive, but it’s the norm for all school-going children. The private tutoring industry is pegged at a whopping $10.8 billion. This is where AI in Education can be revolutionary by creating a level playing field in access to quality education.
India’s AI in education policy has to be multilayered because there is massive wealth and accessibility inequality. The priority should be to help a significant percentage of children from the poorer and rural communities to leap-frog, so that India achieves a revolution in education by creating a pool of skilled and job-ready population in a generation. To achieve this, AI in education must be designed keeping in mind the poor and uneducated and treat it as a digital public good. The first step is to integrate AI across education systems in public or state-funded schools across all states in different languages.
The systems must be inspired and adopt the best practices, such as UNESCO and the OECD recommendations of using AI as a means to enhance cognitive development and lifelong learning, provided that systems remain human-centred, inclusive, and transparent. Political will and public-private partnership are the keys to the success of a project of such magnitude, with the stated mission of AI for good for everyone, everywhere.
One example from India’s context is OpenAI’s Study Mode feature, launched in July 2025 to help students with technical subjects such as maths and computer science. The design was not prompt-dependent, but the idea was to enforce learning behaviour. The key features of the study mode are
OpenAI’s Head of Education, Leah Belsky, explained that Study Mode originated from field observations in India, where families were spending a significant portion of their earnings on private tuitions, thus disadvantaging children from economically weaker sections. OpenAI used India as a design laboratory by beta testing students nationwide, including those preparing for highly competitive medical and engineering exams. Participants were not individually named for privacy, but research indicates that it spanned everyday learning to high-stakes prep, providing feedback that shaped personalisation and scaffolding features. There are no specific cities mentioned, with very few details about the research outcome
Subsequent OpenAI Learning Accelerator (launched August 2025) collaborated with IIT Madras ($500K research on AI learning outcomes), AICTE, Ministry of Education, and ARISE schools, distributing 500K ChatGPT licenses to educators and students.
OpenAI’s Study Mode supports 11 languages with Voice Capability, which is a significant stride for an inclusive education. The voice-enabled interaction can greatly help first-generation students with learning difficulties and those alienated from traditional classrooms. This approach aligns closely with India’s National Education Policy (NEP) 2020, which emphasises foundational literacy, inquiry-based learning, and the reduction of rote memorisation, along with OpenAI’s own policy “to democratise encouragement, guidance, and confidence — especially for learners who lack access to quality teachers or tutors.”
But to address the learning needs of India’s poorest, AI tools alone will not help. India needs a digital infrastructure and an innovative way to fund digital devices, such as a specially designed tablet that can work as a slate. It has to be an interactive device that is given to all students at a subsidised or no cost. India needs an AI-in-Education Code of Conduct — a governance framework developed through multi-stakeholder consultation, including students, teachers, parents, and civil society. This Code should balance personalisation with autonomy, innovation with equity, and data utility with privacy. Without this, we risk deploying AI that works technically but fails socially.
The approach has to be collaborative, and the programme roll-out has to be meticulously planned because this is where most well-intended projects fail. Every tablet distributed has to have insurance. There has to be a soft penalty system, like “access blocked”, that puts the onus of ownership and accountability of the devices on parents. A behaviour change program must run in parallel, along with investments in control and support centres providing 24/ 7 support through chatbots with human oversight. The progress tracker of every child should be available with a unique password and ID, and where parents are uneducated, teachers will assist.
For children coming from poorer families, there has to be an incentive model. If a child does well, credit points could be provided that parents use for redeeming stationery or paying for something else. Such models can arrest dropouts, encourage parents not to withdraw children from school to join the labour market. Foundational learning with AI should help children transition into an AI-driven skill development program that can help them get decent jobs early on. It should lead to creating a credible employability pipeline. Fixing education has ripple effects on other factors, such as child labour, human trafficking of girls, in addition to providing a demographic dividend to the country.
India has a good measurement infrastructure already in place. The Parakh Rastriya Savbekshan, a national assessment system, tested approximately 2.3 million students from 782 districts covering classes 3, 6 and 9, so India already has a baseline data of massive scale, so when AI tools are introduced, the baseline data can be used for comparison to evaluate how AI is making a difference.
The new thrust of India’s National Curriculum Framework shifts focus to building core competencies of what students can do rather than which class they sat in. An AI tool can be used to find out if students can do things they couldn’t before. It will help in creating a competency–based measurement framework for judging AI’s real impact.
India has also created a multi-level assessment dissemination system where data is shared through workshops at national, regional and state levels to inform practical action. This is ideal for infrastructure for AI programs because AI impact measurement needs a similar pipeline to share results. Since India has already built a system, AI evaluation can leverage it.
Accelerating AI in education needs a bold vision, good data and transparent enforcement of the existing mechanisms. The ASER 2024 report, released by Pratham Foundation in January 2025, shows the highest recorded reading levels for Class 3 government school students since the survey began 20 years ago. This is attributed to focused government programs like the NIPUN Bharat Mission.
ASER data has been referenced in 105 parliamentary questions, used by NITI Aayog (India’s planning body), and cited in the World Bank’s World Development Report. For AI in education to succeed. India has proven it can build trusted measurement systems. Now, as AI tools like Study Mode are deployed, these same systems can track whether accessible AI is delivering on its promise — providing personalised, patient support for foundational skills that current interventions cannot fully reach.
From a policy perspective, India’s Governance Policy Architecture, such as the DPDP ACT 2023, has provisions for student data protection, mandatory algorithmic audits for assessment tools following the framework’s fairness requirements, and teacher empowerment over surveillance. Special child safety provisions — explicitly flagged in the Guidelines — should prevent AI systems from exploiting developing minds. Further integration through DIKSHA, Bhashini, and PARAKH offers the infrastructure; the governance framework offers the guardrails. The opportunity is transformative; responsible deployment ensures no child is left behind.
How AI was used in researching and writing his article
AI Claude was used as an augmentation tool while writing this article. Peplexity was used for deep research. Every citation and data was verified. Gemini was used for infographics. The author has also created a Claude project for iterating on editorial flow discussion etc, but the author ensures that the outcome is his own.
AI resilience isn’t about tech companies fixing their systems — it’s about how society adapts when AI safeguards fail. This article explains how the resilience framework, through Avoidance, Defence, and Remedy interventions, prepares society to manage risks from increasingly powerful and accessible AI systems.
On the 6th of December 2025, a seventeen-year-old boy in Japan was arrested for carrying out a cyberattack on the server of Kaikatsu Frontier, an internet café chain operator. The teenager hacked the system by generating code using conversational AI, thereby compromising the data of 7.3 million customers and disrupting all business operations. AI Incident Database reported that the suspect’s prompts to the AI “concealed malicious intent”. The suspect had a case history of unrelated credit card fraud.
The case reveals several facts about the current state of AI. That a teenager was able to access and comprehend powerful AI capabilities easily and cheaply by building sophisticated code, presumably with basic knowledge. Together with the right prompts, he was able to attack a server system that would otherwise need years of expertise to plan and execute.
Clearly, the company’s cyber infrastructure was weak; the AI models are already diffused into our lives. We cannot control who builds it or how people will use it. There will be elements that misuse it, thus exposing the dual capability nature of the models to benefit or harm society.
This pattern isn’t isolated to Japan. In Denmark, a 22-year-old used AI to research how to injure his father without killing him. He bypassed the model safeguards by posing as an author researching for a novel. The AI provided a detailed plan to execute the intended harm.
Let’s look at the magnitude of the teenager’s crime: millions affected by just one action, disrupting business operations all along. How would you protect millions, and in cases where the data exposure results in other harm, such as breaking into the banking systems? How and who would compensate the “millions” of victims?
In the article, I will discuss how adopting the AI resilience framework — which focuses on the societal adaptation through avoidance, defence and remedy interventions — can enhance collective responsibility of Tech companies, governments and society to mitigate AI risks.
AI Resilience complements traditional model-level safety approaches. While the conventional AI safety focuses on training data quality, safeguards, and capability restrictions, resilience addresses what happens when the model-level safety fails or is bypassed, as examples above demonstrate. It is more about how to deal with safety when it is easily accessible (as already is), becoming cheaper to use and deploy for high-end tasks and has a potential for dual use capability. In such a scenario, we need a societal response in terms of adaptation to Advanced AI systems.
Societal Adaptation or AI Resilience begins when AI models are deployed and get diffused. It is at this wide-scale use that new risks are discovered that might slip through even the best of safeguarding measures organically built into the systems, so adaptive interventions are meant to mitigate the harm that might arise in the specific use case of the AI. For example, students use conversational AI for coding or as tutors, but in the above example, one of the teenagers used it for committing a cybercrime.
AI Resilience has a three-part framework:
1. Avoidance: Interventions that stop harmful use before it happens. This includes laws against AI-assisted crimes, age restrictions on accessing harmful content, and monitoring systems that detect suspicious activity. Avoidance aims to prevent attacks from occurring in the first place.
2. Defence — It aims to prevent harm even when the misuse occurs by building systems that withstand it. Interventions include cybersecurity infrastructure that deters cyberattacks, spam filters that catch phishing, and public awareness campaigns to raise awareness of AI harms.
3. Remedy: This intervention reduces the downstream impact after the harm occurs. This includes legal actions such as arrest and prosecution, victim compensation (e.g., people losing money through banking scams) and other rapid responses to contain damage. Remedy faces severe challenges at scale. While legal systems can prosecute one attacker, like in the Denmark case above, they struggle to help millions of victims, like in the Japanese case, where the data of the customers could have been used to hack banking systems.
Let’s deconstruct the Japanese case using the above framework:
Avoidance Failure: In the case of the Japanese teenager, despite his earlier credit card fraud history, the cyber laws didn’t deter him. No monitoring systems tracked his activities, nor any surveillance detected his planning phase, and no age restriction prevented his access to AI’s code-generation capabilities.
Defence Failure: The company’s weak cybersecurity mechanism was vulnerable, and the intrusion succeeded without detection. The generated code succeeded in breaching servers and exposing the data of 7.3 million customers.
Remedy partially worked: Legal accountability led to the arrest of the attacker after months of investigation. However, it compromised the privacy data of 7.3 million people and caused operational havoc which cannot be undone. So one case of misuse of an AI model targeted millions at scale, explaining why remediation systems struggle with AI-enabled mass harm.
Let’s look at another example to explain the Adaptive Cycle. The Economist this month published an article about “How AI is Rewiring Childhood”, signalling exciting opportunities but also cautioning about the ominous risk.
The article talks about the enormous power of AI to transform education by creating a level playing field (if supported by the right policies). A child educated in Hindi medium in remote Bihar might be able to develop cognitive and comprehensive skills like his counterpart educated in English in New Delhi, thus overcoming language barriers — that is a possibility. But there could be a host of other AI tools that could be used for storytelling and learning, but the actual risk lies in the larger societal impact of such tools.
A child begins to show trust and reliance tendencies on the AI system that always speaks in a friendly tone. The AI grasps the child’s psychology and answers in a specific way. The child becomes intolerant of others who disagree with him or are critical of him. Worst, we could soon have a generation of kids who grow up with poor social and networking skills.
The Adaptive cycle puts the onus not just on the company but also on the governments, schools, parents and researchers to create an enabling environment where a child grows up like a human despite the wide-scale AI diffusion. Here’s how:
Identify and Assess Risks (Stage 1): Researchers assess impacts and identify risks. Companies use real-time data for R&D for product enhancement. Schools evaluate the social implications of wide-scale integration of AI in education.
Develop Responses (Stage 2): Governments create industry benchmarks and age-appropriate guidelines. Tech companies put added parental control features, etc. Schools teach social skills and provide, encourage, and promote outdoor games.
Implement Responses (Stage 3): Companies provide explicit literature to parents. Companies conduct surveys with parents. Parents monitor usage and set boundaries. Schools are adopting social skills programs. Government regulations are enforced across the industry.
Loop back to Stage 1 (creating a continuous cycle): New AI capabilities emerge, interventions don’t work as expected, and children develop new use patterns that weren’t previously anticipated, thus making resilience a continuous process.
As LLM model become widespread and easily accessible, the scope of its misuse grows exponentially given the dual use nature of the technology. Mitigating AI risk is not the concern of AI companies alone but a collective action where tech companies, governments, security agencies and the civil society organisations come together by building mechanisms, so that social adaptation and use of AI is secured for the wider benefit of all. The two examples from Japan and Denmark show that non-suspicious and determined actors can bypass safeguards through simple social engineering causing unprecedented harm and havoc.
AI Resilience is our capacity to adapt when AI safeguards fail. The response requires multi-layered interventions across Avoidance, Defence and Remedy so the scale asymmetry of AI misuse can be addressed.
At the societal level, this requires Avoidance through robust cyber laws, professional enforcement agencies that understand AI Governance. Defence through AI products that inform consumers about their potential misuse, awareness amongst parents and teachers about the dangers of over-risk. Remedy through effective enforcement and comprehensive victim support, including psycho-social support.
Governments around the world are building AI Resilience through a combination of laws and oversight. For example, India has adopted a polycentric adaptive model relying on voluntary compliance rather than centralised regulation. It operationalises the adaptive cycle using real-world evidence to inform periodic revisions.
Resilience through continuous adaptation offers a path forward when model-level safety inevitably fails. Tech companies, governments, and society must accelerate the adaptive cycle — building infrastructure to live safely with AI even when its guardrails are breached.
References
Bernardi, J. (2024, August 3). Resilience and adaptation to advanced AI. Achieving AI Resilience. https://achievingairesilience.substack.com/p/resilience-and-adaptation-to-advanced
Bernardi, J., Mukobi, G., Greaves, H., Heim, L., & Anderljung, M. (2024). Societal adaptation to advanced AI. arXiv preprint arXiv:2405.10295. https://arxiv.org/abs/2405.10295
Denmark assault case [Incident 851]. (2025). AI Incident Database. Partnership on AI. https://incidentdatabase.ai/cite/851
How AI is rewiring childhood. (2024, December 7). The Economist. https://www.economist.com/leaders/2024/12/05/how-ai-is-rewiring-childhood
Ministry of Electronics and Information Technology. (2024, November). India AI governance guidelines. Government of India. https://www.meity.gov.in/
Osaka cyberattack [Incident 1047]. (2025, January 18). AI Incident Database. Partnership on AI. https://incidentdatabase.ai/cite/1047
Sahariah, Sutirtha (2025, December 4). Deconstructing India’s AI governance framework. Medium. https://medium.com/@suti011/deconstructing-indias-ai-governance-framework-abd81f6b4cdf
Exploring the intersection of artificial intelligence and policy development in the modern digital landscape.
Comprehensive research on human trafficking, forced labor, and modern slavery across South Asia and beyond.
Developing communication strategies and content for international development organizations.
Helping organizations build and maintain their reputation through strategic messaging and crisis communication.
Expert in designing and conducting qualitative studies, focus groups, and stakeholder interviews.
Published by Routledge – A groundbreaking study on women’s empowerment in Nepal’s informal entertainment sector.
Lessons from the Informal Entertainment Sector in Nepal (2022)
Published by Routledge – A groundbreaking study on women’s empowerment in Nepal’s informal entertainment sector.
This book presents an analysis of the concepts of female empowerment and resilience against violence in the informal entertainment and sex industries.
Generally, the key debates on sex work have centred on arguments proposed by the oppressive and empowerment paradigms. This book moves away from such debates to look widely at the micro issues such as the role of income in the lives of sex workers, the significance of peer organisations and networks of women, and how resilience is enacted and empowerment experienced. It also uses positive deviancy theory as a useful strategy to bring about notable changes in terms of empowerment and agency for women working in this sector and also for addressing the wider issues of migration, HIV/AIDS, and violence against women and girls. The focus is on moving beyond a victimisation framework without downplaying the extent of the violence that women in this industry experience. It conceptualises the theories of empowerment and power which have not been tested against women who work in this sector, combined with in-depth interviews with women working in the industry as well as academics, activists, and personnel in the NGO and donor sector. In doing so, it informs the reader of the numerous social, political, and economic factors that structure and sustain the global growth of the industry and analyses the diverse factors that lead many thousands of women and girls around the world to work in this sector.
The work presents an important contribution to the study of citizenship and rights from a non-Western angle and will be of interest to academics, researchers, and policymakers across human rights, sociology, economics, and development studies.
Knowledge for Change? Lessons from co-developing a research agenda on survivor engagement. November 2023.
Comprehensive review of promising practices across South Asia
a case study from the frontline source area in India
View Research → | Read Policy Impact → | Read Reports from the Project →
Knowledge for Change?
November 2023
Introduction and context ‘Survivor engagement’, understood as the involvement of people with lived experience in policy and programming, has seemingly moved to the centre of efforts to address modern slavery and human trafficking, but how can it really shift the way that these issues are tackled? As practice in this area is underdeveloped, the production of knowledge is likely to be crucial in this, changing approaches and responses through the development of new concepts, interpretations, tools and instruments that can be embedded in policy and practice. This report presents a summary of new findings and reflections from an ongoing and collaborative initiative to develop a research agenda through the lens of survivor engagement. It builds on a project that explored promising practices of lived experience engagement in modern slavery policy and programming and which took place in 2022.1 Researchers at the University of Liverpool, with funding from Foreign, Commonwealth, and Development Office (FCDO), built an international network of researchers and consultants to explore effective methods and practices involving persons with lived experience in modern slavery policy and programming. Recognising the collaborative research’s significance, the network secured additional funding from the Modern Slavery and Human Rights Policy and Evidence Centre (Modern Slavery PEC) to expand their study between March and July 2023. This expansion enabled a deeper exploration of engagement with first-hand experience and expertise in policy and programme systems.
View Research → | Read Policy Impact → | Read Reports from the Project →
Fair purchasing practices in garment supply chains.
connecting theory and practice
Matthew Anderson, Tamsin Bradley, Sutirtha Sahariah
connecting theory and practice
Matthew Anderson, Tamsin Bradley, Sutirtha Sahariah
Abstract
In this chapter, we investigate the experience of Fair Trade organisations and how they have translated Fair Trade principles into practice in their value chains. In particular, we focus on the implementation of responsible purchasing practices related to: Equal Partnership, Collaborative Production Planning and Fair Payment Terms. We argue that, if supported, Fair Trade organisations have the potential to be industry front-runners and demonstrate fair purchasing practices that can be replicated and scaled across the garment sector.
Research Consultation • Strategic Communications • Policy Analysis • Writing & Content Creation • Training & Workshops
Phone: +91 9818965091 Email: suti011@gmail.com