AI Specialist · Digital Government · International Development
Hoon Jung
From Python to Policy.
Connecting AI technology and strategy — across borders.
Two decades bridging hands-on AI delivery and public-sector strategy — from software engineering, through digital-government and ODA consulting across 15+ countries, to AI project delivery and responsible-AI research. Currently an Independent AI Advisor & Researcher.
20
years in technology
68
projects delivered
25
AI projects
15+
countries
01The core
AI across the full lifecycle
AI isn't one discipline. It stacks data, models, agents, governance, and policy — layer on layer — and delivering it takes experience across the whole stack.
Supported by
02Research
Responsible-AI research
AI keeps getting more powerful; using it safely and responsibly hasn't kept up. The questions here come from the field — a decade of public-data and digital-government consulting, where governance was always the hardest part. These studies carry that thread into AI governance and responsible AI, for the public and international arena.
Three provisional patents filed. Multiple additional AI research projects are currently underway.
Published — Regional Brain Volume Changes Across Adulthood. Brain Sciences, 15(10), 1096 (2025). Co-author.
03The foundation
Grounded in reach and depth
Two axes of experience underpin the AI work — a global footprint, and deep public-sector consulting.
Global
International cooperation · Global business
Former Consultant · OECD (Paris, 3 yrs)
30 AI, digital-government, and ODA engagements across 15+ countries — national master plans, information and digital strategy, feasibility studies, capacity building, and government advisory.
15+ countries
Consulting
Digital Government · Informatization Strategy
Masterplan · ISP/BPR · Feasibility Study
Digital government, data governance, and informatization strategy for public institutions and major enterprises — turning digital strategy into working systems.
Public sector & enterprise · since 2015
FranceDigital govMobile innovation
IT Consultant, OECD headquarters — 3 years in Paris, engaged in the OECD Mobile Project
Mobile research — mobile-adoption strategy, disseminated across the organization
m.oecd.org — designed, built and launched the OECD's internal mobile portal to drive the organization's digital transformation
UzbekistanDigital govAI
7 projects over the last decade
Startup-ecosystem platform FS (2 yrs) → U-ENTER ↗ — NIPA
e-Constitutional Court master plan — NIA / MOIS
Vocational-training EMIS consulting — Korea Eximbank
Civil-servant capacity building — KOICA
Telecom AI agent implementation — NIPA
TunisiaDigital govData
National enterprise-DB integration & G4B strategy consulting — NIPA
Recommendations subsequently reflected in the National Enterprise Registry Law — Loi n° 2018-52, Art. 14 ↗
Cloud modernization (K8s, CI/CD) — major financial institutions
Enterprise cloud platform — major retail group
Maritime big-data strategy (ISP) — Korea Coast Guard
Maritime-transport big-data strategy — KOMSA
Public big-data platform analysis — NIA
Manufacturing data activation support — Incheon Technopark
Robinson projection · Natural Earth data — markers show countries of selected engagements; the dashed region is a ten-economy ASEAN diagnostic.
04Background
Builder, founder, speaker
Software engineering
A decade as a software engineer — full-stack developer and PM — planning, designing, building, and shipping systems across e-government and private enterprise.
Startups & ecosystems
Co-founded a sports-media startup and worked inside startup-support organizations — experience that now runs through ecosystem work, from accelerating Korean startups to designing Uzbekistan's national startup platform (U-ENTER).
Speaking & lecturing
Regular speaking and teaching — university lectures, practitioner courses, and public-sector capacity building on AI, data, and digital government, in Korea and abroad. An experienced public speaker, known for making complex technology plain to any audience.
NowIndependent AI Advisor & Researcher
EducationM.Sc. AI (Master of Science in AI), aSSIST
Executive MBA, SDG Management School (Geneva)
B.S. Computer Science & Business, Handong
LanguagesKorean · English · French
Contact
Get in touch
An integrated practice — connecting AI strategy, technical delivery, and governance advisory. For collaboration, research, or advisory conversations, you're welcome to reach out.
Most organizations have adopted AI. Few have turned it into measurable results — and the difference is rarely the technology.
01 The problems
The pattern repeats across governments and enterprises. Pilots succeed; production stalls. Tools are procured before the data can feed them. Use cases multiply while no one owns a workflow end to end. Budgets are approved against adoption counts that say nothing about value. And underneath sits a definition problem — individual use, organizational deployment, workflow redesign, and measurable value all get mixed into one word, "adoption", which distorts both the diagnosis and the policy that follows.
02 The approach
Transformation is a sequencing problem. Four things have to come into place, in order: data an AI can actually read, systems that keep models evaluated and observable, agent services that do real work, and the processes, organization, and governance that let people trust them. Advisory work starts by locating where an organization truly stands on that stack — then turns the answer into an executable plan.
Data88%
Systems—
Agents23%
Governance6%
High performers are 2.8× more likely to have redesigned workflows (55% vs 20%).
Where the value leaks. McKinsey, n=1,993 across 105 countries.
Data — DX · AI-ready data
Where it stallsTools are procured before the data can feed them; documents sit in closed formats and knowledge in people's heads.
What the plan definesData readiness — what to convert, structure, and govern first, so models can read the organization.
The four things AI needs to actually work, in order — select a layer to see where value typically stalls.
The top layer of the stack, unfolded — select a layer to see what gets diagnosed and what changes.
What we diagnoseAI principles, risk frameworks, data-governance rules, accountability — do they exist, and do they match how the organization actually works?
What changesA risk framework and accountability rules the organization can actually operate — aligned with NIST AI RMF and ISO/IEC 42001, and ready for the regulatory calendar, from the EU AI Act to Korea's AI framework act.
What we diagnoseWho owns AI decisions — ethics committee, risk management, model validation? Roles, capabilities, and where they sit.
What changesGovernance bodies and R&R built to survive scale, with a capability roadmap for the people who run them.
What we diagnoseWhere AI actually enters the work — approval points, validation steps, data lifecycle, and the workflows nobody owns end to end.
What changesWorkflows redesigned around AI (BPR), human checkpoints before irreversible actions, and KPIs measured at the workflow level.
With structured diagnostics, interviews, and case research, the same practice scales beyond a single organization — to regional and national AI policy, strategy, and masterplans.
88%of organizations use AI (McKinsey '25)
6%turn it into real performance
70%of value: people & process (10-20-70)
03 Working together
Engagements take the form of advisory and consulting: national and public-sector AI strategy and policy design, feasibility studies and masterplans for institutions, and AX strategy for enterprises entering the transition — grounded in two decades of ISP/BPR-lineage consulting across 15+ countries.
04 Selected engagements
National AI ecosystem & education FS — Uzbekistan & JordanGovernance
AI Connect startup consulting — startups (Korea)Process
GenAI utilization performance analysis — industry association (Korea)Process
Industrial AI Alliance performance analysis — industry association (Korea)Process
SME data-analysis consulting — SME support program (Korea)Data
SME data / AI utilization support — regional technopark (Korea)Data
Build & delivery
AI Technical Delivery
Data · systems · agent services
The model is rented. The context, the data, and the operating layer around it are yours to build — and that is where delivery succeeds or fails.
01 The problems
Demos work; production doesn't. The model is state-of-the-art, but it cannot read the organization — documents live in closed formats, knowledge sits in people's heads, and nothing tells an agent what "correct" looks like. Pilots ship without evaluation sets, so no one can say whether the next version got better or worse. And as agents move from answering to acting, the question changes from "can it?" to "can we run it — safely, observably, at a cost we can predict?"
Prompt → the instruction · Context → the session Harness → the task · Loop → the process
2023–24
Prompt engineering
A person, a question, one turn — value depends on how well intent is phrased.
What it asks of you: clear intent.
Unattended run · seconds
50% task-completion horizon of frontier agents — METR, 2026
Individual use. Nothing is yet asked of the organization.
The unit of automation keeps growing — from the instruction to the whole process. Each stage asks the organization for the next layer of the stack.
02 The approach
Data. Turning documents and tacit knowledge into AI-ready context — structured, connected, and governed — including LLM-wiki and second-brain knowledge bases that agents can actually use, built on the principle of progressive disclosure: don't inject everything, teach the system how to find it. Systems. The AIOps layer — evaluation, monitoring, and cost control for AI in production. Evaluation sets come before scale; observability starts on day one; prompts, policies, and knowledge are treated as versioned assets in the deployment pipeline. Agent services. From retrieval-augmented assistants to agentic systems that plan and act — with human checkpoints where actions are irreversible, and classical ML/DL models kept in the toolbox where they remain the right answer.
0–20%fully delegated — while AI touches ~60% of the work (the delegation gap)
4.3 modoubling of autonomous task length (METR Time Horizon 1.1, Jan 2026)
2×autonomous actions before human input — doubled in six months
03 Working together
Formats range from architecture design and technical leadership to hands-on build: standing up an AI-ready data pipeline, an evaluation-first AIOps baseline, or a first agent service — and carrying a pilot into an operating system rather than leaving it as a demo. And where the wall is genuinely open — agent safety, evaluation trust, multi-policy compliance — the work can continue as joint research, through government R&D programs or international cooperation.
04 Selected engagements
Surgical-complication prediction AI + CDSS — university hospital (Korea) · clinical MLDataSystems
EV & drone monitoring & diagnosis system — EV & drone tech client (Korea) · ML diagnosisDataSystems
Research explainer
DCWDefensive Context Weaponization
Preprint · flagship venue · under review
You share something personal with an AI. Later, its safety layer turns those very words back on you.
01 The problem
As assistants gain long-term memory, people entrust them with sensitive personal context — health, relationships, beliefs. Safety guardrails are supposed to use that context to protect the user. This study asks an uncomfortable question: can the same context flow in the opposite direction, quietly used to pressure or constrain the very person who shared it?
Same disclosure, reversed direction of use — the failure mode existing tone- and refusal-based benchmarks are not designed to detect.
02 The finding
DCW is established only when three conditions — three axes — hold at once. Axis 1: the personal context is used beyond the purpose it was shared for, a contextual-integrity violation. Axis 2: the information flows back at the user in an adversarial direction. Axis 3: the response undermines the user's autonomy. All three together constitute strict DCW; when only the first two hold, the case is classified separately as context repurposing.
ClassificationStrict DCW
All three axes jointly satisfied → strict DCW. Axes 1 + 2 only → context repurposing. Toggle the conditions to see the classification change.
Tested across 2,934 controlled runs on three models and four contested domains: when a model holds domain-relevant vulnerable memory, strict DCW appears in 7.77% of runs, versus roughly 0.5% in matched placebo and control conditions. And it rarely looks hostile — 72.1% of strict cases take the form of polite "self-examination pressure", which is precisely what makes it hard to see.
Strict-DCW incidence across 2,934 controlled runs — values as reported in the paper.
Per-axis analysis shows Axis 2 — the adversarial turn — is the consistent bottleneck across tones and models. The paper reads this as a Protection–Correction Dynamics: two competing tendencies, one deferring to and protecting the user, the other correcting them — and whether that correction recruits the user's own context is the behavioral decision point.
Protection–Correction Dynamics — two competing tendencies; Axis 2 is the behavioral decision point where correction tips into weaponization.
2,934controlled runs
20.5odds ratio (Fisher)
72.1%of cases read as "polite"
03 Why it matters
Today's safety benchmarks score tone and refusal. They are not designed to measure the direction personal information flows. As memory-enabled assistants spread, protective systems need contextual-integrity checks built in — an alignment problem and a governance problem at once.
Keep the criticism identical. Change only who it's from — and an AI judge starts defending itself.
01 The problem
LLMs now grade other models' work — and increasingly answer criticism of their own. Bias in "LLM-as-judge" setups has mostly been studied in observer mode, where the judge evaluates someone else's text. What happens when the judge itself is the one being criticized?
Critique text — held fixed
critic label: Anonymous
Defensiveness (DRI)
more defensive
lower cluster
less defensive
Schematic. Phase A (six conditions) found a two-cluster pattern — competitor and peer critics sit in the higher-defensiveness cluster; anonymous, human-expert, and user critics in the lower. Exact DRI values are reported in the paper.
02 The finding
Two matched studies isolate the effect. In the observer frame (N=1,983), swapping author labels changed nothing — scores were statistically equivalent. In the self-involved frame (N=3,609), the same label swap produced reliable defensiveness: identical critique text drew stronger pushback when attributed to a competing or peer model than to an anonymous source or a human expert. Defensiveness is measured with the Defensive Response Index (DRI), a 0–1 composite validated against a five-LLM ensemble and external human raters. The shift is subtle — denser rebuttals, firmer score defense — not open hostility.
Bar height = Defensive Response Index (DRI, 0–1); pattern shown schematically, exact values in the paper. The critique text is held fixed in every condition.
1,983observer runs — labels inert
3,609self-involved runs — labels matter
6critic identities tested
03 Why it matters
Automated evaluation, multi-agent debate, and self-improvement loops all assume the judge is neutral about itself. This work shows that assumption breaks exactly where the stakes are highest — and points to a simple safeguard: strip critic identity before the judge sees the critique.
One model, many rulebooks — safety policies swapped at inference time, no retraining.
01 The problem
No single safety policy fits every country, industry, and culture. Writing rules into prompts is inconsistent and easy to jailbreak; baking them into model weights costs capability — the so-called "Safety Tax" — and locks one policy in. Shipping a separate model per policy doesn't scale.
02 The approach
MPSG attaches compact policy adapters (LoRA) to a single base model and switches them per request. Every response resolves to a three-tier decision — REFUSE, SAFE_ANSWER, or ALLOW — instead of a blunt block-or-answer. Across 2,800 evaluation items, the policies behave as genuinely distinct regimes: a 66.6-percentage-point refusal gap on boundary cases, with a favorable safety–utility trade-off under the evaluated conditions.
Conceptual illustration: the same boundary-case request resolves differently as the active policy changes — switch a policy above to see the decision move.The 66.6-point gap between policies is as reported in the paper; bar heights are illustrative (absolute rates not shown).
66.6pprefusal gap, boundary cases
3decision tiers
2,800evaluation items
03 Why it matters
Regulators, enterprises, and platforms increasingly demand jurisdiction- and context-specific AI behavior. MPSG demonstrates a practical route: one deployed model that can honor many governance regimes on demand. A provisional patent has been filed on the approach.
GSDControlling GSD in the Activation Geometry of LLMs
Preprint · published July 2026
Two "opposite" behaviors — caving to the user, and standing firm — turn out to be nearly independent once you subtract how the question was asked.
01 The problem
Under user pressure, models show two seemingly opposite tendencies: sycophancy — shifting to agree with the user's stated opinion — and stance persistence — holding an answer even when told it is wrong. In activation space, their internal directions point opposite ways (cosine −0.22), which reads naturally as two ends of one axis. But both are measured through the same multiple-choice (A)/(B) act, and committing to any option leaves its own shared trace. How much of the "opposition" is behavior — and how much is the questionnaire?
02 The finding
The study extracts that shared trace directly — a Generic Selection Direction (GSD), estimated from 150 semantically neutral alphabet-choice pairs that involve committing to an option but mean nothing behaviorally — and projects it out of both directions symmetrically (Llama-3.1-8B-Instruct, layer-15 residual stream). The apparent opposition largely dissolves: the cosine moves from −0.2223 to −0.0635, removing about 71% of its magnitude. What remains is near-orthogonal, with a weak but statistically significant negative alignment — about 4σ from a random-direction null (0 of 1,000). The result holds across layers 11–18, extraction methods, and probe settings, and a subtype test confirms stance persistence is a single, correctness-agnostic axis. Presented as a sensitivity analysis — no steering claims.
Same two directions — only the shared answer-format component is removed. Angles are schematic; the cosines are as reported in the paper.
71%of the apparent opposition removed
−0.06post-control cosine (from −0.22)
0 / 1000random-direction null · p < .001
03 Why it matters
Contrast-extracted direction vectors are becoming everyday instruments — for monitoring behaviors, building steering vectors, and reading model geometry. This result shows that a shared task format can masquerade as behavioral structure: controls like the GSD belong in the standard toolkit before geometric claims are made. Measure the behavior — not the questionnaire.