Invest1 distinct publisher2 min readPublished
The AI Frontier Division built Cofa-Probe on OpenAI's Codex to score four stages of human-agent collaboration, which tells you roughly what Krafton now thinks a finished portfolio piece is worth as hiring evidence.
The Investor · Invest desk

Compiled by The InvestorSomething wrong?How this is made
The load-bearing detail is the evidence Krafton keeps. Its tool captures source code, conversation logs, prompts, tool settings and traces of verification runs from a candidate's session [5], which is five classes of evidence where the conventional screen kept one, the submitted artifact [3][1]. Four stages get scored instead: problem definition and prioritization, the quality of AI collaboration and instructions, evidence-gathering and decision-making, and iterative improvement and error recovery [4], four scored process stages set against the single graded deliverable a conventional screen produced [4].
Krafton is candid about the weak joint. With an LLM, the same submission can return different results from one run to the next [7], so evaluation criteria and score conversion are controlled by code while the judgments that require reading context go to LLM agents [9], and a meta harness repeatedly tests and refines the grading agent's execution structure [8]. Fencing the scoring arithmetic in code is the right call, and it also means the score travels no further than the harness that produced it. The instrument itself was built using OpenAI's Codex [2], so the rig that measures how well you direct agents is itself a directed-agent product.
What Krafton is not doing with that engineering time is the part I would price. The AI Frontier Division built an internal hiring instrument roughly ten months after the October declaration that the company would become AI-first [10][2], and then took it outside the building, grading tasks set by CJ Olive Young at a hiring-linked hackathon at PUBG Seongsu in Seongdong [12], with plans to widen the tool's scope [15] and an ambition, in AI Frontier Division head Park Jae-min's words, to set a new standard for assessing AI-native talent [13]. The premise a Krafton official gives is that problem-solving ability does not show up in the output alone [14].
There are a few ways this plays differently. The process signal is real, and the hires ramp faster than portfolio-screened ones did. Or the four named categories become coachable, and candidates learn to narrate hypotheses at the grader rather than form them [4]. Or the rig is the deliverable [15][13], and hiring quality is a by-product for whoever eventually runs it on their own applicants. This is probably wrong, but I would weight the second, on the grounds that a published rubric decays into performance and this one's stages are already in print [4]. What would move me is dull evidence: retention or ramp numbers on Cofa-Probe hires against the previous cohort, or a variance figure showing two runs of one submission landing on the same score [7].
Ranked by verification strength, evidence, and original report placement.
Krafton (259960) has developed its own AI-based hiring and assessment solution, Cofa-Probe, built by its AI Frontier Division, and has begun using it in recruiting for developer positions, according to information technology industry sources on the 29th.
Cofa-Probe is a large language model-based solution developed using OpenAI's Codex, and it evaluates how candidates define problems and verify results while working with AI.
Cofa-Probe creates a virtual environment resembling the actual work of an AI engineer, in which it assesses how candidates identify and solve problems while interacting with multiple task agents; unlike conventional assessments that look only at the finished product, it structures the entire way a candidate works with AI.
The solution breaks the collaboration process into four categories: problem definition and prioritization, quality of AI collaboration and instructions, evidence-gathering and decision-making, and iterative improvement and error recovery.
During assessment, Cofa-Probe automatically collects not only source code but also conversation logs, prompts, tool settings and traces of verification runs.
After a task ends, candidates take part in a follow-up question-and-answer session explaining their output and decision-making process, and the system generates a reference report for interviewers.
Distinct publishers with included, body-backed reporting in this cluster.
en.sedaily.com
1 article · August 28, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
product
OpenAI's plan to hand everyone a coding agent leaves the hard part to the model1 distinct publisher
build
Twenty-three security checks, zero coverage: AI coding agents as build-pipeline attack surface1 distinct publisher
product
Four leaderboards, four denominators: what you buy when you standardize on a coding agent1 distinct publisher
build
Codex can now ask and keep going, which deletes the only checkpoint you were getting for free1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One telling, company-shaped
Everything here — the Codex build, the four scored categories, the capture list, the harness — reaches the reader through a single trade report that credits unnamed industry sources and two Krafton speakers. The mechanism detail is specific enough to be checkable, which is a point in its favour, but nobody has checked it, and the one claim that would need data behind it (that the meta harness makes grading reproducible) arrives with no measurement at all.
Two live uses, no scale
This is past the demo stage: the tool is described as already grading developer applicants, and it graded real submissions at a hackathon run with an outside partner in CJ Olive Young. What is missing is any sense of size — how many candidates, which teams, what share of hiring — so the footprint reads as two confirmed instances rather than a programme.
Mechanism grounded, ambition floating
Mildly overstated, and the gap sits in a specific place. The plumbing is described soberly — Krafton even names the failure mode, that an LLM grades the same submission differently on different runs. Then the fix for that failure mode is asserted, and the division head reaches for "a new standard for assessing AI-native talent" on the strength of one hackathon and an unquantified recruiting pilot.
Builder, user and narrator are one party
Krafton built the tool, uses the tool, and is the only party describing it — while ten months into a publicly declared AI-first pivot that this story conveniently evidences. The Cofathon is itself recruiting-brand marketing with a consumer partner attached. None of that makes the account wrong; it does mean every favourable detail originates with the side that benefits from the pivot looking real.
Trust the description, not the claims about it
We can be fairly confident the system exists and works roughly as described — the detail is too granular and too internally consistent to be vapour, and it comes from the people who built it. Confidence drops sharply on anything about how well it works: reproducibility, fairness, hiring outcomes and future reach all rest on Krafton's word in a single outlet.