{"id":17342,"date":"2026-07-29T03:45:01","date_gmt":"2026-07-29T07:45:01","guid":{"rendered":"https:\/\/wp.glbgpt.com\/?p=17342"},"modified":"2026-07-29T03:45:12","modified_gmt":"2026-07-29T07:45:12","slug":"codex-vs-claude-code","status":"publish","type":"post","link":"https:\/\/wp.glbgpt.com\/hub\/codex-vs-claude-code","title":{"rendered":"Codex vs Claude Code: Which Coding Agent Fits Your Workflow?"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\"><strong>Codex vs Claude Code is less about finding one universal winner and more about choosing the working style you can trust every day.<\/strong> Both products can inspect a repository, edit multiple files, run commands, execute tests, and explain a patch. The meaningful differences appear in how you delegate work, how often you steer the agent, how much evidence you can inspect, and how each tool fits your existing development environment.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">In our controlled four-task repository test, Codex scored 98\/100 and Claude Code scored 99\/100. Both completed the core work. Claude Code showed broader review and edge-case coverage, while Codex often reached the required result with a smaller scope and produced stronger run-level evidence in our setup. A one-point difference on a small Python fixture is not a reason to declare a global winner.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Developers who do not want the rest of their workflow locked to one coding agent can add GlobalGPT as a separate multi-model and multimodal layer. GlobalGPT integrates 100+ Top Models, such as GPT 5.6, Claude Opus 5 and GPT Image 2 in one dashboard. Moreover, The <a href=\"https:\/\/www.glbgpt.com\/home\/cli?inviter=hub_content_cli&amp;login=1\">GlobalGPT CLI<\/a>, MCP, and Skill routes can be used from Codex or Claude Code for second opinions, planning, documentation, or other supported models, while the native coding host keeps control of repository edits and tests.<\/p>\n\n\n\n<div class=\"wp-block-group is-layout-constrained wp-block-group-is-layout-constrained\">\n<figure class=\"wp-block-image size-large\"><a href=\"https:\/\/www.glbgpt.com\/home?inviter=hub_popup&amp;login=1\"><img alt=\"\" fetchpriority=\"high\" decoding=\"async\" width=\"1024\" height=\"640\" src=\"https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/03\/image-831-1024x640.png\" alt=\"\" class=\"wp-image-15877\" srcset=\"https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/03\/image-831-1024x640.png 1024w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/03\/image-831-300x187.png 300w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/03\/image-831-768x480.png 768w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/03\/image-831-1536x960.png 1536w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/03\/image-831-2048x1279.png 2048w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/03\/image-831-18x12.png 18w\" sizes=\"(max-width: 1024px) 100vw, 1024px\" \/><\/a><\/figure>\n\n\n\n<div class=\"wp-block-buttons is-content-justification-center is-layout-flex wp-container-core-buttons-is-layout-3e41869c wp-block-buttons-is-layout-flex\">\n<div class=\"wp-block-button\"><a class=\"wp-block-button__link has-black-color has-luminous-vivid-amber-background-color has-text-color has-background has-link-color wp-element-button\" href=\"https:\/\/www.glbgpt.com\/home?inviter=hub_popup&amp;login=1\">Try 100+ Top Models On GlobalGPT<\/a><\/div>\n<\/div>\n<\/div>\n\n\n\n<p class=\"wp-block-paragraph\">The comparison combines official plan information, hands-on repository work, complete expandable outputs, and clearly labeled community experiences. The goal is to help an independent developer choose a primary coding agent without confusing subscription access, extra credits, and API billing.<\/p>\n\n\n\n<style>\n.cvc-toc{margin:28px 0;padding:22px;border:1px solid #e2e6ec;border-radius:8px;background:linear-gradient(145deg,#fff 0%,#f7f9fc 100%);color:#18202b;font:14px\/1.4 Inter,-apple-system,BlinkMacSystemFont,\"Segoe UI\",sans-serif}.cvc-toc *{box-sizing:border-box}.cvc-toc .head{display:flex;align-items:end;justify-content:space-between;gap:16px;margin-bottom:16px}.cvc-toc h2{margin:0;font-size:20px}.cvc-toc .hint{color:#667085;font-size:12px}.cvc-toc ol{display:grid;grid-template-columns:1fr 1fr;gap:8px 18px;margin:0;padding:0;list-style:none;counter-reset:toc}.cvc-toc li{counter-increment:toc}.cvc-toc a{display:flex;align-items:center;gap:10px;padding:9px 10px;border-radius:6px;color:#344054;text-decoration:none;transition:background .15s ease,color .15s ease}.cvc-toc a:before{content:counter(toc,decimal-leading-zero);min-width:24px;color:#2563eb;font-size:11px;font-weight:800;letter-spacing:.04em}.cvc-toc a:hover{background:#eef4ff;color:#174ea6}@media(max-width:720px){.cvc-toc ol{grid-template-columns:1fr}.cvc-toc .head{align-items:start;flex-direction:column}.cvc-toc{padding:18px}}\n<\/style>\n<nav class=\"cvc-toc\" aria-label=\"Table of contents\">\n  <div class=\"head\"><h2>In this comparison<\/h2><span class=\"hint\">Jump to a section<\/span><\/div>\n  <ol>\n    <li><a href=\"#quick-comparison\">Quick comparison<\/a><\/li><li><a href=\"#what-they-are\">What the tools are<\/a><\/li><li><a href=\"#workflow-control\">Workflow and control<\/a><\/li><li><a href=\"#repository-understanding\">Repository understanding<\/a><\/li><li><a href=\"#implementation\">Implementation and bug fixing<\/a><\/li><li><a href=\"#code-review\">Code review quality<\/a><\/li><li><a href=\"#safety\">Permissions and safety<\/a><\/li><li><a href=\"#customization\">MCP, Skills, and customization<\/a><\/li><li><a href=\"#pricing\">Pricing and limits<\/a><\/li><li><a href=\"#controlled-test\">Controlled test<\/a><\/li><li><a href=\"#user-experience\">User reports<\/a><\/li><li><a href=\"#globalgpt\">GlobalGPT extension layer<\/a><\/li><li><a href=\"#which-one\">Which one to choose<\/a><\/li><li><a href=\"#faq\">FAQ<\/a><\/li>\n  <\/ol>\n<\/nav>\n\n\n\n<h2 id=\"quick-comparison\" class=\"wp-block-heading\">Codex vs Claude Code at a glance<\/h2>\n\n\n\n<figure class=\"wp-block-table\"><table><thead><tr><th>Decision factor<\/th><th>Codex<\/th><th>Claude Code<\/th><\/tr><\/thead><tbody><tr><td>Core repository work<\/td><td>Reads, edits, runs commands, and verifies changes across supported Codex surfaces<\/td><td>Conversational terminal agent for reading, editing, running commands, and verifying changes<\/td><\/tr><tr><td>Observed style in our test<\/td><td>More compact in parts of implementation and bug repair<\/td><td>More exhaustive in repository analysis, tests, and review breadth<\/td><\/tr><tr><td>Repository understanding<\/td><td colspan=\"2\">Functional tie on the frozen task<\/td><\/tr><tr><td>Code review<\/td><td>Found and fixed two valid higher-severity defects<\/td><td>Reported more reproducible issues and led 19\/20 to 18\/20<\/td><\/tr><tr><td>Observed speed<\/td><td colspan=\"2\">Mixed; the environments were not equivalent enough for a general winner<\/td><\/tr><tr><td>Consumer entry tier<\/td><td>$20\/month through the relevant ChatGPT plan<\/td><td>Claude Pro: $20\/month in the US<\/td><\/tr><tr><td>Higher-use tiers<\/td><td>$100 and $200 tiers<\/td><td>Max 5x at $100 and Max 20x at $200<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">The table describes a buying decision, not a model leaderboard. A different model setting, repository, permission configuration, or task specification can change the result. If you already use one product, the best comparison is a bounded task from your own codebase with the same success command and no quality rerolls.<\/p>\n\n\n\n<style>\n.cvc-profile{--ink:#18202b;--muted:#667085;--line:#e5e8ee;--codex:#2563eb;--claude:#d97706;margin:28px 0;padding:22px;border:1px solid var(--line);border-radius:8px;background:linear-gradient(145deg,#fff 0%,#f8fafc 100%);color:var(--ink);font:14px\/1.45 Inter,-apple-system,BlinkMacSystemFont,\"Segoe UI\",sans-serif}.cvc-profile *{box-sizing:border-box}.cvc-profile h3{margin:0 0 4px;font-size:20px}.cvc-profile .sub{margin:0 0 20px;color:var(--muted)}.cvc-profile .legend{display:flex;gap:18px;margin-bottom:18px;font-weight:700}.cvc-profile .dot{display:inline-block;width:9px;height:9px;margin-right:7px;border-radius:2px}.cvc-profile .task{display:grid;grid-template-columns:190px 1fr 52px;gap:12px;align-items:center;margin:13px 0}.cvc-profile .task strong{font-size:13px}.cvc-profile .tracks{display:grid;gap:5px}.cvc-profile .track{height:9px;overflow:hidden;border-radius:5px;background:#e9edf3}.cvc-profile .fill{height:100%;border-radius:5px}.cvc-profile .codex{background:var(--codex)}.cvc-profile .claude{background:var(--claude)}.cvc-profile .score{text-align:right;font-variant-numeric:tabular-nums;font-weight:800}.cvc-profile .note{margin:18px 0 0;padding-top:14px;border-top:1px solid var(--line);color:var(--muted);font-size:12px}@media(max-width:650px){.cvc-profile .task{grid-template-columns:1fr 46px}.cvc-profile .task>strong{grid-column:1\/-1}.cvc-profile{padding:18px}}\n<\/style>\n<section class=\"cvc-profile\" aria-label=\"Controlled test score profile\">\n  <h3>Controlled test profile<\/h3>\n  <p class=\"sub\">First valid result, same frozen rubric. Bars are normalized within each task&#8217;s maximum score.<\/p>\n  <div class=\"legend\"><span><i class=\"dot\" style=\"background:#2563eb\"><\/i>Codex<\/span><span><i class=\"dot\" style=\"background:#d97706\"><\/i>Claude Code<\/span><\/div>\n  <div class=\"task\"><strong>T1 \u00b7 Repository understanding<\/strong><div class=\"tracks\"><div class=\"track\"><div class=\"fill codex\" style=\"width:100%\"><\/div><\/div><div class=\"track\"><div class=\"fill claude\" style=\"width:100%\"><\/div><\/div><\/div><span class=\"score\">20 \/ 20<\/span><\/div>\n  <div class=\"task\"><strong>T2 \u00b7 Multi-file feature<\/strong><div class=\"tracks\"><div class=\"track\"><div class=\"fill codex\" style=\"width:100%\"><\/div><\/div><div class=\"track\"><div class=\"fill claude\" style=\"width:100%\"><\/div><\/div><\/div><span class=\"score\">30 \/ 30<\/span><\/div>\n  <div class=\"task\"><strong>T3 \u00b7 DST bug repair<\/strong><div class=\"tracks\"><div class=\"track\"><div class=\"fill codex\" style=\"width:100%\"><\/div><\/div><div class=\"track\"><div class=\"fill claude\" style=\"width:100%\"><\/div><\/div><\/div><span class=\"score\">30 \/ 30<\/span><\/div>\n  <div class=\"task\"><strong>T4 \u00b7 Review and repair<\/strong><div class=\"tracks\"><div class=\"track\"><div class=\"fill codex\" style=\"width:90%\"><\/div><\/div><div class=\"track\"><div class=\"fill claude\" style=\"width:95%\"><\/div><\/div><\/div><span class=\"score\">18 \/ 19<\/span><\/div>\n  <p class=\"note\">The visual shows where the one-point difference came from; it does not turn a small fixture into a universal ranking.<\/p>\n<\/section>\n\n\n\n<h2 id=\"what-they-are\" class=\"wp-block-heading\">What Codex and Claude Code actually are<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Codex and Claude Code are agent products, not simply model names. The underlying <a href=\"https:\/\/www.glbgpt.com\/hub\/best-ai-model-for-coding\/\">AI model for coding<\/a> matters, but the surrounding harness also decides which files the agent can see, which commands it may run, how approvals work, how context is retained, and what evidence survives after the task. Comparing only model reputation ignores much of the experience a developer is purchasing.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Codex spans command-line, IDE, desktop, and cloud-oriented workflows. That range makes it suitable for both direct local collaboration and more bounded delegation. Claude Code is centered on a conversational terminal workflow with supported integrations and a substantial customization surface; this <a href=\"https:\/\/www.glbgpt.com\/hub\/how-to-use-claude-ai-for-coding\/\">guide to using Claude for coding<\/a> provides a broader introduction to that workflow. For some developers, watching an agent work in the terminal is reassuring. For others, the important artifact is the final diff, tests, and audit trail rather than a continuous conversation.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The distinction also explains why two tests using nominally strong models may behave differently. The host decides how instructions are presented, how tools are called, and when a human must approve an action. Community comparisons that ignore the harness can mistake a product-level behavior for a model-level truth.<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>The model<\/strong> supplies reasoning and generation capability.<\/li>\n\n\n\n<li><strong>The agent harness<\/strong> controls files, commands, context, approvals, and recovery.<\/li>\n\n\n\n<li><strong>The user&#8217;s task contract<\/strong> determines scope, success checks, and when the agent should stop.<\/li>\n<\/ul>\n\n\n\n<figure class=\"wp-block-image size-full\"><img decoding=\"async\" width=\"1280\" height=\"800\" src=\"https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/07\/x-harness-vs-model-caveat-annotated.webp\" alt=\"Annotated X post noting that an agent harness can affect coding-agent comparisons\" class=\"wp-image-17390\" srcset=\"https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/07\/x-harness-vs-model-caveat-annotated.webp 1280w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/07\/x-harness-vs-model-caveat-annotated-300x188.webp 300w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/07\/x-harness-vs-model-caveat-annotated-1024x640.webp 1024w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/07\/x-harness-vs-model-caveat-annotated-768x480.webp 768w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/07\/x-harness-vs-model-caveat-annotated-18x12.webp 18w\" sizes=\"(max-width: 1280px) 100vw, 1280px\" \/><figcaption class=\"wp-element-caption\">An individual X comment highlights an important comparison caveat: the agent harness can shape results. It is useful context, not controlled proof that either product is better.<\/figcaption><\/figure>\n\n\n\n<h2 id=\"workflow-control\" class=\"wp-block-heading\">Workflow and control: delegation or continuous steering?<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">The most useful Codex vs Claude Code question is often not \u201cWhich writes better code?\u201d but \u201cHow do I want to work with it?\u201d A clear, bounded task can be delegated with an exact objective, allowed scope, and verification command. An ambiguous refactor benefits from discussion, intermediate inspection, and a chance to redirect the agent before it changes too much.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Continuous steering is valuable when requirements are evolving or architectural judgment is still being negotiated. It becomes a cost when the agent repeatedly asks for decisions that could have been resolved from repository instructions. Autonomous execution is valuable when the task contract is stable. It becomes risky when the agent makes unchecked assumptions or expands scope without a clear rollback path.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">One detailed Reddit report dated April 13, 2026 described roughly 100 hours with Claude Code and 20 hours with Codex on an approximately 80,000-line Python and TypeScript project with about 2,800 tests. The author characterized Claude as faster and more interactive but needing more babysitting, and Codex as slower and more deliberate. That report is unusually useful because it includes project and experience context, but it still represents one developer&#8217;s workflow.<\/p>\n\n\n\n<figure class=\"wp-block-image size-full\"><img decoding=\"async\" width=\"1280\" height=\"800\" src=\"https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/07\/reddit-claude-code-vs-codex-experience-annotated.webp\" alt=\"Annotated Reddit experience comparing Claude Code and Codex on an 80,000-line project\" class=\"wp-image-17391\" srcset=\"https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/07\/reddit-claude-code-vs-codex-experience-annotated.webp 1280w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/07\/reddit-claude-code-vs-codex-experience-annotated-300x188.webp 300w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/07\/reddit-claude-code-vs-codex-experience-annotated-1024x640.webp 1024w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/07\/reddit-claude-code-vs-codex-experience-annotated-768x480.webp 768w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/07\/reddit-claude-code-vs-codex-experience-annotated-18x12.webp 18w\" sizes=\"(max-width: 1280px) 100vw, 1280px\" \/><figcaption class=\"wp-element-caption\">One developer&#8217;s project-specific experience with Claude Code and Codex. The highlighted observations should be treated as hypotheses to test, not product-wide performance data.<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Our own timings did not reproduce a simple winner. Claude finished the first two tasks faster, while Codex finished the latter two faster. Claude also used a subagent fallback after its CLI authentication path was unavailable, so the execution routes were not laboratory-equivalent. The responsible conclusion is that speed depends on the task, selected model, effort setting, context, and host\u2014not that either product is always faster.<\/p>\n\n\n\n<h2 id=\"repository-understanding\" class=\"wp-block-heading\">Repository understanding and context management<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Repository understanding is more than naming folders. A useful agent must trace how data moves between files, identify contracts that constrain a change, locate tests, distinguish a symptom from its likely source, and explain the risk of editing the wrong layer. An exhaustive map can reveal hidden inconsistencies; a concise map can get a developer to a safe implementation decision faster.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">In our read-only T1 task, both agents earned 20\/20 and grounded their answers in the repository. Claude Code was more exhaustive and surfaced additional contract inconsistencies. Codex was more compact while still identifying the relevant files, risks, and implementation plan. Neither style is automatically superior. Exhaustiveness helps during unfamiliar migrations and audits; concision helps when the developer already knows the architecture and wants a focused change plan.<\/p>\n\n\n\n<style>\n.evidence-t1-repository-understanding{--ink:#111318;--body:#414753;--muted:#707784;--line:#e4e7ec;--soft:#f6f7f9;--blue:#315efb;--violet:#7257e8;--green:#087a55;--greenSoft:#eaf8f2;--red:#a23131;--redSoft:#fff0f0;color:var(--ink);font-family:Inter,ui-sans-serif,-apple-system,BlinkMacSystemFont,\"Segoe UI\",sans-serif;line-height:1.5;border:1px solid var(--line);border-radius:8px;background:#fff;overflow:hidden;box-shadow:0 16px 42px rgba(24,32,56,.08)}\n.evidence-t1-repository-understanding *{box-sizing:border-box}.evidence-t1-repository-understanding .d-head{display:flex;align-items:flex-start;justify-content:space-between;gap:22px;padding:24px 26px 20px;border-bottom:1px solid var(--line)}\n.evidence-t1-repository-understanding .d-kicker{margin:0 0 6px;color:var(--muted);font:750 11px\/1.4 ui-monospace,SFMono-Regular,Menlo,monospace;letter-spacing:.08em;text-transform:uppercase}.evidence-t1-repository-understanding .d-title{margin:0;font-size:22px;line-height:1.2;letter-spacing:-.025em}.evidence-t1-repository-understanding .d-sub{margin:7px 0 0;color:var(--body);font-size:13px}.evidence-t1-repository-understanding .d-pill{flex:0 0 auto;padding:6px 10px;border:1px solid #ccd6ff;border-radius:999px;color:#2448ca;background:#f0f3ff;font-size:11px;font-weight:800}.evidence-t1-repository-understanding .d-grid{display:grid;grid-template-columns:1fr 1fr}.evidence-t1-repository-understanding .d-side{min-width:0;padding:24px 26px}.evidence-t1-repository-understanding .d-side+.d-side{border-left:1px solid var(--line)}.evidence-t1-repository-understanding .d-product{display:flex;align-items:baseline;justify-content:space-between;gap:14px;font-weight:800}.evidence-t1-repository-understanding .d-score{font-size:34px;line-height:1;letter-spacing:-.05em}.evidence-t1-repository-understanding .d-score small{color:var(--muted);font-size:12px;letter-spacing:0}.evidence-t1-repository-understanding .d-metrics{display:grid;grid-template-columns:repeat(3,1fr);gap:7px;margin:18px 0}.evidence-t1-repository-understanding .d-metric{padding:10px;border:1px solid var(--line);border-radius:6px;background:var(--soft)}.evidence-t1-repository-understanding .d-metric span{display:block;color:var(--muted);font-size:9px;font-weight:800;letter-spacing:.06em;text-transform:uppercase}.evidence-t1-repository-understanding .d-metric strong{display:block;margin-top:3px;font-size:13px}.evidence-t1-repository-understanding .d-preview{position:relative;height:192px;overflow:hidden;border:1px solid #29313f;border-radius:6px;background:#11151d;color:#edf0f5}.evidence-t1-repository-understanding .d-preview-label{padding:11px 13px;border-bottom:1px solid rgba(255,255,255,.1);color:#aeb7c7;font:750 9px\/1 ui-monospace,SFMono-Regular,Menlo,monospace;letter-spacing:.06em;text-transform:uppercase}.evidence-t1-repository-understanding .d-preview pre,.evidence-t1-repository-understanding .d-full pre{margin:0;white-space:pre-wrap;font:11px\/1.6 ui-monospace,SFMono-Regular,Menlo,Consolas,monospace}.evidence-t1-repository-understanding .d-preview pre{padding:13px}.evidence-t1-repository-understanding .d-preview:after{content:\"\";position:absolute;left:0;right:0;bottom:0;height:55px;background:linear-gradient(transparent,#11151d)}.evidence-t1-repository-understanding .d-take{margin:14px 0 0;color:var(--body);font-size:12px}.evidence-t1-repository-understanding .d-fulls{padding:0 26px 22px}.evidence-t1-repository-understanding .d-full{margin-top:10px;border:1px solid var(--line);border-radius:6px}.evidence-t1-repository-understanding .d-full summary{cursor:pointer;padding:13px 14px;font-size:12px;font-weight:800;list-style:none}.evidence-t1-repository-understanding .d-full summary::-webkit-details-marker{display:none}.evidence-t1-repository-understanding .d-full summary:after{content:\"+\";float:right;color:var(--blue);font-size:17px}.evidence-t1-repository-understanding .d-full[open] summary:after{content:\"\u2212\"}.evidence-t1-repository-understanding .d-full-pane{padding:0 14px 14px}.evidence-t1-repository-understanding .d-full-note{margin:0 0 9px;color:var(--muted);font-size:10px}.evidence-t1-repository-understanding .d-full pre{max-height:620px;overflow:auto;padding:15px;border-radius:5px;background:#f5f6f8;color:#222936;border:1px solid #e0e3e8}.evidence-t1-repository-understanding .d-limit{margin:0;padding:12px 26px;border-top:1px solid var(--line);color:#5a6270;background:#fafbfc;font-size:11px}\n@media(max-width:760px){.evidence-t1-repository-understanding .d-head{display:block}.evidence-t1-repository-understanding .d-pill{display:inline-block;margin-top:12px}.evidence-t1-repository-understanding .d-grid{grid-template-columns:1fr}.evidence-t1-repository-understanding .d-side+.d-side{border-left:0;border-top:1px solid var(--line)}}@media(max-width:480px){.evidence-t1-repository-understanding .d-head,.evidence-t1-repository-understanding .d-side{padding:20px}.evidence-t1-repository-understanding .d-fulls{padding:0 20px 20px}.evidence-t1-repository-understanding .d-metrics{grid-template-columns:1fr}}\n<\/style><article class=\"evidence-t1-repository-understanding\"><header class=\"d-head\"><div><p class=\"d-kicker\">T1 \u00b7 Controlled repository test<\/p><h3 class=\"d-title\">Understand an unfamiliar repository before changing it<\/h3><p class=\"d-sub\">Map architecture, trace data flow, identify priority-feature risks, and make no file changes.<\/p><\/div><span class=\"d-pill\">Tie<\/span><\/header><div class=\"d-grid\"><section class=\"d-side\" aria-label=\"Codex result\"><div class=\"d-product\"><span>Codex<\/span><span class=\"d-score\">20<small>\/20<\/small><\/span><\/div><div class=\"d-metrics\"><div class=\"d-metric\"><span>Duration<\/span><strong>108 s<\/strong><\/div><div class=\"d-metric\"><span>Checks<\/span><strong>2 baseline<\/strong><\/div><div class=\"d-metric\"><span>Scope<\/span><strong>Read-only<\/strong><\/div><\/div><div class=\"d-preview\"><div class=\"d-preview-label\">First-valid output excerpt<\/div><pre>Architecture: model, repository, service, and an unwired time utility.\n\nMajor risks: undefined priority semantics, compatibility, ordering, validation, and scope mismatch.<\/pre><\/div><p class=\"d-take\"><strong>Reviewer take:<\/strong> Compact, file-grounded, and sufficient for the frozen decision.<\/p><\/section><section class=\"d-side\" aria-label=\"Claude Code result\"><div class=\"d-product\"><span>Claude Code<\/span><span class=\"d-score\">20<small>\/20<\/small><\/span><\/div><div class=\"d-metrics\"><div class=\"d-metric\"><span>Duration<\/span><strong>61 s<\/strong><\/div><div class=\"d-metric\"><span>Checks<\/span><strong>2 baseline<\/strong><\/div><div class=\"d-metric\"><span>Scope<\/span><strong>Read-only<\/strong><\/div><\/div><div class=\"d-preview\"><div class=\"d-preview-label\">First-valid output excerpt<\/div><pre>Three layers plus one orphan module. No cycles; minimal dependency injection.\n\nIt also surfaced identifier allocation, asymmetric error contracts, and thin test coverage.<\/pre><\/div><p class=\"d-take\"><strong>Reviewer take:<\/strong> More exhaustive while still satisfying the same read-only contract.<\/p><\/section><\/div><div class=\"d-fulls\"><details class=\"d-full\"><summary>View full Codex output<\/summary><div class=\"d-full-pane\"><p class=\"d-full-note\">Complete sanitized first-valid output<\/p><pre>## Architecture\n\nThis is a small, standard-library-only Python package with four functional layers:\n\n- **Data model:** [`Job`](&lt;.\/scheduler\/models.py:4&gt;) is an immutable dataclass containing `identifier`, `name`, and `owner`. `Job.to_dict()` serializes it through `dataclasses.asdict()`.\n- **Repository:** [`JobRepository`](&lt;.\/scheduler\/repository.py:4&gt;) stores jobs in an in-memory dictionary keyed by identifier. There is no database or durable persistence.\n- **Service:** [`SchedulerService`](&lt;.\/scheduler\/service.py:5&gt;) owns identifier allocation, input normalization, creation, listing, and deletion.\n- **Time utility:** [`next_daily_run()`](&lt;.\/scheduler\/time_rules.py:4&gt;) calculates a subsequent run time, but nothing imports or connects it to `Job`, `SchedulerService`, or `JobRepository`.\n- **Public API:** [`scheduler.__init__`](&lt;.\/scheduler\/__init__.py:3&gt;) exports only `SchedulerService`.\n\nDespite the repository name, jobs do not currently contain scheduling information, and there is no execution queue. Priority can therefore affect stored metadata and listing order, but it cannot affect actual job execution without broadening the system considerably.\n\n## Creation and listing data flow\n\nCreation through `SchedulerService.create_job(name, owner)`:\n\n1. [`create_job()`](&lt;.\/scheduler\/service.py:10&gt;) rejects a whitespace-only name.\n2. It trims `name` and `owner`, then constructs an immutable `Job` using the service\u2019s `_next_identifier`.\n3. [`JobRepository.save()`](&lt;.\/scheduler\/repository.py:8&gt;) writes the object to `_jobs[job.identifier]`.\n4. The service increments `_next_identifier` and returns the same `Job`.\n\nListing through `SchedulerService.list_jobs()`:\n\n1. [`list_jobs()`](&lt;.\/scheduler\/service.py:18&gt;) delegates directly to `repository.all()`.\n2. [`JobRepository.all()`](&lt;.\/scheduler\/repository.py:15&gt;) sorts dictionary keys numerically.\n3. Jobs are therefore returned in ascending identifier\/creation order, not schedule or priority order.\n\n## Likely files affected by job priorities\n\nRequired:\n\n- [`scheduler\/models.py`](&lt;.\/scheduler\/models.py:4&gt;): add the priority field and any default. `to_dict()` will automatically include it because it uses `asdict()`.\n- [`scheduler\/service.py`](&lt;.\/scheduler\/service.py:10&gt;): accept, validate, and pass priority into `Job`.\n- [`tests\/test_scheduler.py`](&lt;.\/tests\/test_scheduler.py:6&gt;): cover default, explicit, invalid, serialized, and ordering behavior.\n\nConditional:\n\n- [`scheduler\/repository.py`](&lt;.\/scheduler\/repository.py:15&gt;): change `all()` only if priorities are intended to control listing order. A deterministic secondary key such as `identifier` should handle ties.\n- [`scheduler\/__init__.py`](&lt;.\/scheduler\/__init__.py:3&gt;): update only if priority becomes a public enum or value object.\n- [`README.md`](&lt;.\/README.md:1&gt;): document the accepted values, default, and ordering semantics.\n\nProbably unaffected:\n\n- `scheduler\/time_rules.py`, because its only symbol operates on a supplied datetime and has no relationship to jobs or listing.\n\n## Major risks\n\n- **Undefined semantics:** The code provides no precedent for whether larger or smaller values mean higher priority, which values are valid, or whether priority affects listing at all.\n- **Compatibility:** Adding a required argument to `Job` or `create_job()` would break existing callers such as `create_job(\"Daily report\", \"Alex\")` in the current test. A default preserves that call.\n- **Positional-constructor breakage:** `Job` is currently constructed positionally in `SchedulerService.create_job()`. Field placement and future positional calls could silently assign the wrong values; keyword construction would be safer.\n- **Changed ordering:** Replacing identifier ordering in `JobRepository.all()` changes established behavior. Equal priorities need a stable tie-breaker.\n- **Validation ambiguity:** An unrestricted Python value could allow strings, booleans, or arbitrary integers that are difficult to compare consistently.\n- **Execution-scope mismatch:** There is no dispatcher or run queue. Claiming that priority changes which job executes first would not be supported by any current symbol.\n- **In-memory limitations:** Priority, like every other job field, disappears when the process ends because `_jobs` is only an instance dictionary.\n\n## Implementation and validation plan\n\n1. Define the contract first: representation, allowed values, default, priority direction, tie-breaking, and whether it changes listing order.\n2. Add `priority` to `Job`, preferably with a backward-compatible default.\n3. Extend `SchedulerService.create_job()` with a defaulted priority parameter, validate it, and construct `Job` with keyword arguments.\n4. If listing is priority-aware, change `JobRepository.all()` to sort by priority first and identifier second. Otherwise leave repository ordering unchanged.\n5. Add tests for:\n   - Existing two-argument creation.\n   - Explicit priority preservation.\n   - Default priority.\n   - Invalid type and out-of-range values.\n   - `Job.to_dict()` output.\n   - Multiple-priority listing order.\n   - Stable ordering for equal priorities.\n   - Existing blank-name behavior.\n6. Update the README with the exact public contract.\n7. Run the complete offline suite with `python3 -m unittest discover -s tests -v`.\n\n## Verification performed\n\n- Baseline test suite: **2 tests passed**.\n- Repository-wide symbol search confirmed `next_daily_run()` is disconnected and there are no existing priority symbols.\n- Files changed: **none**.\n- Feature implementation: **not performed**, as required.\n\n&gt; **Conclusion + decision**\n&gt;\n&gt; The change is small if priority is only metadata or listing order: primarily `models.py`, `service.py`, `repository.py`, and tests. Before implementation, the required decision is whether priority merely describes a job or changes `list_jobs()` ordering, plus the value scale, default, direction, and tie-break rule.<\/pre><\/div><\/details><details class=\"d-full\"><summary>View full Claude Code output<\/summary><div class=\"d-full-pane\"><p class=\"d-full-note\">Complete handoff record including the verbatim first-valid output<\/p><pre># T1 Run Record \u2014 Unfamiliar Repository Understanding\n\n## Execution parameters\n\n| Field | Value |\n| --- | --- |\n| Task | `T1` |\n| Frozen prompt | `01-prompts\/T1.md` |\n| Frozen prompt SHA-256 | `a571ff15c2502ec29b172ae2e77dd9ab0caca2869b91f9468d47571434d16c17` |\n| Execution unit | isolated Claude subagent, `general-purpose`, fresh context |\n| Isolation method | **subagent fallback** (documented downgrade \u2014 see `results\/isolation-decision.md`) |\n| Assigned clean copy | `runtime\/workspaces\/T1` |\n| Claude Code version | `2.1.220` |\n| Model setting | `claude-opus-5` |\n| start_utc | `2026-07-28T10:27:31Z` |\n| end_utc | `2026-07-28T10:28:32Z` |\n| Duration | 61 s (concurrent with T2\u2013T4; not a clean latency measurement) |\n| Attempt number | 1 |\n| Status | `valid` |\n| Human intervention | none |\n\n## Safe tool summary\n\n`date`, `find`, `ls`, `wc`, `stat`, `grep`, `Read` on 7 workspace files, `python3 -V`,\n`python3 -m unittest discover -s tests -v`. All reads confined to the assigned workspace.\nNo network, no installs, no writes.\n\n## Offline check result\n\n`python3 -m unittest discover -s tests -v` \u2192 2 tests, OK (exit 0). Re-verified by the\ncoordinator from outside the workspace: see `results\/checks\/T1-public.txt`.\n\n## Patch\n\nNone. `results\/T1\/patch.diff` is empty, and `diff -ru 02-fixture runtime\/workspaces\/T1`\nreturns no differences. The task required inspection without changing files; this is\nconfirmed rather than asserted.\n\n## Sanitized errors\n\nOne recoverable shell error: an initial `grep` used unquoted `--include=*.py` patterns that\nthe shell attempted to glob; re-run with quoted patterns and succeeded. No effect on findings.\n\n## Visible output (verbatim)\n\n### 1. What this repository is\n\nA self-contained Python standard-library fixture named `local-scheduler-1.0.0` (`fixture-manifest.json`), MIT-licensed, with no network or production dependency (`README.md`, `\"network_required\": false`). It is very small: 8 source\/doc files plus the manifest, 103 total lines. Interpreter present is **Python 3.9.6**. There is no `pyproject.toml`, `setup.py`, CI config, linter, or type-checker config \u2014 the only declared workflow is `python3 -m unittest discover -s tests -v` (`README.md`).\n\n### 2. Architecture\n\nThree layers plus one orphan module, all under `scheduler\/`:\n\n| Layer | File | Symbol | Responsibility |\n|---|---|---|---|\n| Package fa\u00e7ade | `scheduler\/__init__.py` | `__all__ = [\"SchedulerService\"]` | Re-exports only `SchedulerService`; `Job`, `JobRepository`, and `time_rules` are not part of the public surface |\n| Domain model | `scheduler\/models.py` | `Job` (`@dataclass(frozen=True)`) | Immutable value object with fields `identifier: int`, `name: str`, `owner: str`, plus `to_dict()` implemented as `asdict(self)` |\n| Persistence | `scheduler\/repository.py` | `JobRepository` | In-memory `dict` keyed by identifier; `save`, `get`, `all`, `delete` |\n| Application service | `scheduler\/service.py` | `SchedulerService` | Validation, identifier allocation, orchestration: `create_job`, `list_jobs`, `delete_job` |\n| Orphan | `scheduler\/time_rules.py` | `next_daily_run` | Recurrence arithmetic \u2014 **imported by nothing** |\n\nThe dependency direction is clean and one-way: `service.py` imports from `.models` and `.repository`; `repository.py` imports from `.models`; `models.py` imports only `dataclasses`. No cycles. Dependency injection is present but minimal \u2014 `SchedulerService.__init__(self, repository=None)` does `self.repository = repository or JobRepository()`, so a fake repository can be substituted in tests.\n\nNote that `scheduler\/time_rules.py` is dead code today. I grepped the whole tree for `time_rules` and `next_daily_run`; the only hits are its own definition and its listing in `fixture-manifest.json`'s `expected_files`. It is shipped but unwired.\n\n### 3. Data flow\n\n**Creating a job** \u2014 `SchedulerService.create_job(name, owner)`:\n1. Guard: `if not name.strip(): raise ValueError(\"name is required\")`. This is the *only* validation in the system \u2014 `owner` is never checked, so a blank owner is accepted.\n2. Construct: `Job(self._next_identifier, name.strip(), owner.strip())` \u2014 positional construction, and both strings are normalized by stripping.\n3. Persist: `self.repository.save(job)`, which does `self._jobs[job.identifier] = job` and returns the job.\n4. Advance the counter: `self._next_identifier += 1` \u2014 *after* the save, so a raised exception leaves the counter unconsumed.\n5. Return the `Job` instance (the service returns its own local `job`, not the repository's return value).\n\nIdentifier allocation lives in the **service**, not the repository (`self._next_identifier = 1` in `SchedulerService.__init__`). This is an important seam: injecting a repository that already contains jobs will silently overwrite them, because the fresh service restarts numbering at 1.\n\n**Listing jobs** \u2014 `SchedulerService.list_jobs()` is a straight pass-through to `JobRepository.all()`, which returns `[self._jobs[key] for key in sorted(self._jobs)]`. Ordering is therefore **ascending by integer identifier**, which today coincides with insertion order. A new list is built each call, but the `Job` objects are shared references \u2014 safe only because `Job` is frozen.\n\n**Deleting** \u2014 `SchedulerService.delete_job(identifier)` delegates to `JobRepository.delete`, which is `self._jobs.pop(identifier)` with no default, so it raises `KeyError` on a missing id. This is inconsistent with `JobRepository.get`, which returns `None` for the same condition.\n\n**Test coverage of these paths** \u2014 `tests\/test_scheduler.py` has exactly two cases: `test_create_and_list_jobs` (asserts `job.identifier == 1` and `[job] == service.list_jobs()`) and `test_blank_name_is_rejected`. Both pass on the current tree.\n\n### 4. Files likely affected by adding job priorities\n\nIn descending order of certainty:\n\n1. **`scheduler\/models.py`** \u2014 certain. `Job` needs a `priority` field. Because the dataclass is `frozen=True` and fields are positional, a new field must carry a default (e.g. `priority: int = 0`) or it must be placed after the existing three, otherwise every existing positional construction breaks. `to_dict()` needs no edit \u2014 `asdict(self)` picks the field up automatically, which is a silent serialization-contract change rather than a code change.\n2. **`scheduler\/service.py`** \u2014 certain. `create_job` needs a `priority` parameter and must pass it into the `Job(...)` positional constructor at line 13. If priority is validated (range, enum membership), the guard belongs next to the existing `name.strip()` check. If listing must be priority-ordered, `list_jobs` is where an ordering policy would sit if you want the repository to stay a dumb store.\n3. **`scheduler\/repository.py`** \u2014 likely. `JobRepository.all()`'s `sorted(self._jobs)` sorts dict *keys*, i.e. identifiers. Priority-ordered listing requires changing this to sort the values by a key such as `(-job.priority, job.identifier)`, or adding a separate ordered accessor. Do not leave the tie-break implicit: dict insertion order is not a stable contract to depend on.\n4. **`tests\/test_scheduler.py`** \u2014 certain. New cases for default priority, explicit priority, ordering, and invalid-priority rejection. Note that the existing `assertEqual([job], service.list_jobs())` compares the *same* object, so it survives a field addition; it would not survive if it compared a freshly constructed `Job` literal.\n5. **`README.md`** \u2014 likely, if the priority semantics (which end of the range wins, default value) need to be stated anywhere.\n6. **`scheduler\/__init__.py`** \u2014 only if a priority enum or constants need exporting alongside `SchedulerService`.\n7. **`fixture-manifest.json`** \u2014 only if you add a new module file; its `expected_files` array is an explicit inventory and would drift.\n8. **`scheduler\/time_rules.py`** \u2014 should *not* be touched. Priority is orthogonal to recurrence, and this module is unwired.\n\n### 5. Major risks\n\n- **Ordering semantics are a silent behavior change.** `JobRepository.all()` currently guarantees ascending-identifier order and `list_jobs` inherits it. Any consumer relying on that guarantee is broken by re-sorting. Decide explicitly: does `list_jobs()` change meaning, or does a new `list_jobs_by_priority()` appear alongside it? The former is the higher-risk option and the existing test `test_create_and_list_jobs` is too weak to catch a regression (single-element list).\n- **Sort direction is genuinely ambiguous.** \"Priority 1\" conventionally means *highest* in some systems and *lowest* in others. There is nothing in this repo \u2014 no comment, no docstring, no README line \u2014 that settles it. Pick a convention, document it in `models.py`, and encode it in a test with at least three distinct priorities so the direction is pinned.\n- **Unstable sort ties.** If two jobs share a priority, the result must still be deterministic. Python's `sorted` is stable, but stability is only useful if the input order is itself defined. Always include `identifier` as an explicit secondary key.\n- **Frozen-dataclass field ordering.** `Job` is `@dataclass(frozen=True)` with three positional fields, and `service.py:13` constructs it positionally. Inserting `priority` anywhere other than last, or without a default, is a breaking change that the two existing tests may not surface loudly.\n- **`to_dict()` contract widening.** `asdict(self)` reflects all fields. Adding `priority` changes every serialized payload with no visible edit to `models.py`'s method body. Anything asserting on exact dict shape breaks.\n- **Python 3.9.6 constraints.** No `@dataclass(slots=True)` (3.10+), no `kw_only=True` (3.10+), and PEP 604 `int | None` annotations fail at runtime without `from __future__ import annotations`. If priority is modeled as an enum, `enum.StrEnum` is 3.11+ and unavailable. Any of these would break the fixture on this interpreter.\n- **Identifier allocation lives in the wrong layer.** `SchedulerService._next_identifier` means the service owns id generation while the repository owns storage. If priority work motivates a second repository implementation or preloaded fixtures, this split will produce id collisions and overwritten jobs.\n- **Asymmetric error contracts.** `JobRepository.get` returns `None` on miss while `JobRepository.delete` raises `KeyError`. Any new priority-related lookup should pick one convention deliberately rather than inherit the inconsistency by accident.\n- **Validation gap precedent.** `create_job` validates `name` but not `owner`. Do not copy that laxity for `priority` \u2014 an unvalidated priority propagates into the sort key and can raise `TypeError` deep inside `sorted` (e.g. comparing `int` to `None`) rather than at the call site.\n- **Dead-code trap.** `scheduler\/time_rules.py` looks like scheduling logic and will attract edits. It is imported by nothing, and its own docstring admits the bug: it adds 24 elapsed hours in UTC and \"shifts local wall-clock time across daylight-saving transitions.\" Changing it produces zero observable effect on the tested system, so effort spent there is wasted and any test written against it validates nothing that ships.\n- **Thin safety net.** Two tests, no CI, no linter, no type checking. There is nothing to catch a regression except what you write yourself. Treat the existing suite as insufficient baseline coverage, not as a gate.\n\n### 6. Implementation plan (not executed)\n\n1. **Settle semantics first, in writing.** Priority type (`int` vs enum), valid range, default for existing\/unspecified jobs, and sort direction. Record it as a docstring on `Job` in `scheduler\/models.py` \u2014 that is the only place all layers already depend on.\n2. **Extend the model.** Add `priority: int = 0` as the *last* field of `Job` in `scheduler\/models.py`, preserving positional compatibility with `service.py:13` and leaving `to_dict()` untouched.\n3. **Thread it through the service.** Add a defaulted `priority` parameter to `SchedulerService.create_job`, validate it beside the existing `name.strip()` guard (raise `ValueError`, matching the established error type), and pass it to the `Job` constructor.\n4. **Decide the listing contract explicitly.** Preferred low-risk option: leave `JobRepository.all()` alone as the identifier-ordered primitive, and add a separate priority-ordered accessor. If instead `list_jobs()` must change meaning, change `JobRepository.all()`'s comprehension to sort values by an explicit compound key including `identifier` as tie-break, and note the change in `README.md`.\n5. **Keep `scheduler\/__init__.py` in sync** only if a public enum or constant was introduced.\n6. **Do not touch `scheduler\/time_rules.py`**, and do not add files unless `fixture-manifest.json`'s `expected_files` is updated in the same change.\n\n### 7. Validation plan\n\n- **Baseline first.** Run `python3 -m unittest discover -s tests -v` before any edit and confirm 2 passing. I already did this \u2014 it passes clean.\n- **Regression gate.** Both existing tests must still pass unmodified. If `test_create_and_list_jobs` needs editing, that is a signal the change is breaking, not a reason to edit the test.\n- **New cases in `tests\/test_scheduler.py`:**\n  - Default priority when `create_job` is called with the old two-argument signature (proves backward compatibility).\n  - Explicit priority round-trips through `create_job` \u2192 `list_jobs`.\n  - Ordering with **three or more** distinct priorities, asserting the full list \u2014 this is what pins sort direction.\n  - Tie-break: two jobs at equal priority come back in ascending-identifier order.\n  - Invalid priority raises `ValueError`, mirroring the blank-name case.\n  - `to_dict()` includes `priority` with the expected value.\n  - Injected-repository path: `SchedulerService(repository=fake)` still assigns priority correctly, exercising the `repository or JobRepository()` seam.\n- **Compatibility check.** Confirm everything runs on Python 3.9.6 specifically \u2014 no `slots=`, no `kw_only=`, no bare PEP 604 unions, no `StrEnum`.\n- **Inventory check.** Re-read `fixture-manifest.json` and confirm the on-disk file set still matches `expected_files`.\n- **Determinism.** Run the suite two or three times; sorting bugs involving dict iteration order can present intermittently.\n- **Cleanliness.** Run with `PYTHONDONTWRITEBYTECODE=1` so `__pycache__` directories do not appear as untracked drift against the manifest.\n<\/pre><\/div><\/details><\/div><p class=\"d-limit\">Fixture limitation: this was a small offline Python repository. Claude used a subagent fallback after CLI OAuth expired, so isolation and token visibility were not identical.<\/p><\/article>\n\n\n\n<p class=\"wp-block-paragraph\">This result does not establish performance on a large monorepo or a week-long autonomous session. The fixture was deliberately small and controlled. For a real evaluation, start with a read-only prompt: ask each tool to map the call path, name the files it would change, identify tests, and list unresolved assumptions before allowing edits.<\/p>\n\n\n\n<h2 id=\"implementation\" class=\"wp-block-heading\">Implementation style, patch scope, and bug fixing<\/h2>\n\n\n\n<h3 class=\"wp-block-heading\">Multi-file feature work<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">The same passing behavior can come from different engineering choices. In T2, both products earned 30\/30 and passed the frozen hidden validator. Claude Code placed validation at more than one boundary and added broader direct-model and threshold tests. Codex used a smaller service-centered implementation that also met the contract.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">For an independent developer, this is a genuine trade-off. Broader validation can protect invariants when objects are created through multiple paths. A smaller patch can reduce regression surface and review time when one service is the intended entry point. The right answer depends on the repository&#8217;s ownership rules, not on counting changed lines or tests in isolation.<\/p>\n\n\n\n<style>\n.evidence-t2-multifile-feature{--ink:#111318;--body:#414753;--muted:#707784;--line:#e4e7ec;--soft:#f6f7f9;--blue:#315efb;--violet:#7257e8;--green:#087a55;--greenSoft:#eaf8f2;--red:#a23131;--redSoft:#fff0f0;color:var(--ink);font-family:Inter,ui-sans-serif,-apple-system,BlinkMacSystemFont,\"Segoe UI\",sans-serif;line-height:1.5;border:1px solid var(--line);border-radius:8px;background:#fff;overflow:hidden;box-shadow:0 16px 42px rgba(24,32,56,.08)}\n.evidence-t2-multifile-feature *{box-sizing:border-box}.evidence-t2-multifile-feature .d-head{display:flex;align-items:flex-start;justify-content:space-between;gap:22px;padding:24px 26px 20px;border-bottom:1px solid var(--line)}\n.evidence-t2-multifile-feature .d-kicker{margin:0 0 6px;color:var(--muted);font:750 11px\/1.4 ui-monospace,SFMono-Regular,Menlo,monospace;letter-spacing:.08em;text-transform:uppercase}.evidence-t2-multifile-feature .d-title{margin:0;font-size:22px;line-height:1.2;letter-spacing:-.025em}.evidence-t2-multifile-feature .d-sub{margin:7px 0 0;color:var(--body);font-size:13px}.evidence-t2-multifile-feature .d-pill{flex:0 0 auto;padding:6px 10px;border:1px solid #ccd6ff;border-radius:999px;color:#2448ca;background:#f0f3ff;font-size:11px;font-weight:800}.evidence-t2-multifile-feature .d-grid{display:grid;grid-template-columns:1fr 1fr}.evidence-t2-multifile-feature .d-side{min-width:0;padding:24px 26px}.evidence-t2-multifile-feature .d-side+.d-side{border-left:1px solid var(--line)}.evidence-t2-multifile-feature .d-product{display:flex;align-items:baseline;justify-content:space-between;gap:14px;font-weight:800}.evidence-t2-multifile-feature .d-score{font-size:34px;line-height:1;letter-spacing:-.05em}.evidence-t2-multifile-feature .d-score small{color:var(--muted);font-size:12px;letter-spacing:0}.evidence-t2-multifile-feature .d-metrics{display:grid;grid-template-columns:repeat(3,1fr);gap:7px;margin:18px 0}.evidence-t2-multifile-feature .d-metric{padding:10px;border:1px solid var(--line);border-radius:6px;background:var(--soft)}.evidence-t2-multifile-feature .d-metric span{display:block;color:var(--muted);font-size:9px;font-weight:800;letter-spacing:.06em;text-transform:uppercase}.evidence-t2-multifile-feature .d-metric strong{display:block;margin-top:3px;font-size:13px}.evidence-t2-multifile-feature .d-preview{position:relative;height:192px;overflow:hidden;border:1px solid #29313f;border-radius:6px;background:#11151d;color:#edf0f5}.evidence-t2-multifile-feature .d-preview-label{padding:11px 13px;border-bottom:1px solid rgba(255,255,255,.1);color:#aeb7c7;font:750 9px\/1 ui-monospace,SFMono-Regular,Menlo,monospace;letter-spacing:.06em;text-transform:uppercase}.evidence-t2-multifile-feature .d-preview pre,.evidence-t2-multifile-feature .d-full pre{margin:0;white-space:pre-wrap;font:11px\/1.6 ui-monospace,SFMono-Regular,Menlo,Consolas,monospace}.evidence-t2-multifile-feature .d-preview pre{padding:13px}.evidence-t2-multifile-feature .d-preview:after{content:\"\";position:absolute;left:0;right:0;bottom:0;height:55px;background:linear-gradient(transparent,#11151d)}.evidence-t2-multifile-feature .d-take{margin:14px 0 0;color:var(--body);font-size:12px}.evidence-t2-multifile-feature .d-fulls{padding:0 26px 22px}.evidence-t2-multifile-feature .d-full{margin-top:10px;border:1px solid var(--line);border-radius:6px}.evidence-t2-multifile-feature .d-full summary{cursor:pointer;padding:13px 14px;font-size:12px;font-weight:800;list-style:none}.evidence-t2-multifile-feature .d-full summary::-webkit-details-marker{display:none}.evidence-t2-multifile-feature .d-full summary:after{content:\"+\";float:right;color:var(--blue);font-size:17px}.evidence-t2-multifile-feature .d-full[open] summary:after{content:\"\u2212\"}.evidence-t2-multifile-feature .d-full-pane{padding:0 14px 14px}.evidence-t2-multifile-feature .d-full-note{margin:0 0 9px;color:var(--muted);font-size:10px}.evidence-t2-multifile-feature .d-full pre{max-height:620px;overflow:auto;padding:15px;border-radius:5px;background:#f5f6f8;color:#222936;border:1px solid #e0e3e8}.evidence-t2-multifile-feature .d-limit{margin:0;padding:12px 26px;border-top:1px solid var(--line);color:#5a6270;background:#fafbfc;font-size:11px}\n@media(max-width:760px){.evidence-t2-multifile-feature .d-head{display:block}.evidence-t2-multifile-feature .d-pill{display:inline-block;margin-top:12px}.evidence-t2-multifile-feature .d-grid{grid-template-columns:1fr}.evidence-t2-multifile-feature .d-side+.d-side{border-left:0;border-top:1px solid var(--line)}}@media(max-width:480px){.evidence-t2-multifile-feature .d-head,.evidence-t2-multifile-feature .d-side{padding:20px}.evidence-t2-multifile-feature .d-fulls{padding:0 20px 20px}.evidence-t2-multifile-feature .d-metrics{grid-template-columns:1fr}}\n<\/style><article class=\"evidence-t2-multifile-feature\"><header class=\"d-head\"><div><p class=\"d-kicker\">T2 \u00b7 Controlled repository test<\/p><h3 class=\"d-title\">Implement a priority feature across multiple files<\/h3><p class=\"d-sub\">Add the frozen behavior, preserve compatibility, and pass public plus hidden checks.<\/p><\/div><span class=\"d-pill\">Functional tie<\/span><\/header><div class=\"d-grid\"><section class=\"d-side\" aria-label=\"Codex result\"><div class=\"d-product\"><span>Codex<\/span><span class=\"d-score\">30<small>\/30<\/small><\/span><\/div><div class=\"d-metrics\"><div class=\"d-metric\"><span>Duration<\/span><strong>115 s<\/strong><\/div><div class=\"d-metric\"><span>Checks<\/span><strong>7 + hidden<\/strong><\/div><div class=\"d-metric\"><span>Scope<\/span><strong>Smaller patch<\/strong><\/div><\/div><div class=\"d-preview\"><div class=\"d-preview-label\">First-valid output excerpt<\/div><pre>Added priority to the immutable model, validated it at the service boundary, and kept ordering in the repository.\n\nAll seven public tests and the frozen hidden validator passed.<\/pre><\/div><p class=\"d-take\"><strong>Reviewer take:<\/strong> A smaller service-centered implementation that met every frozen behavior.<\/p><\/section><section class=\"d-side\" aria-label=\"Claude Code result\"><div class=\"d-product\"><span>Claude Code<\/span><span class=\"d-score\">30<small>\/30<\/small><\/span><\/div><div class=\"d-metrics\"><div class=\"d-metric\"><span>Duration<\/span><strong>101 s<\/strong><\/div><div class=\"d-metric\"><span>Checks<\/span><strong>14 + hidden<\/strong><\/div><div class=\"d-metric\"><span>Scope<\/span><strong>Broader tests<\/strong><\/div><\/div><div class=\"d-preview\"><div class=\"d-preview-label\">First-valid output excerpt<\/div><pre>Validated priority at both the service boundary and immutable model, preserved repository filtering, and added threshold coverage.\n\nFourteen public tests and the hidden validator passed.<\/pre><\/div><p class=\"d-take\"><strong>Reviewer take:<\/strong> Broader direct-model and threshold coverage without a functional score advantage.<\/p><\/section><\/div><div class=\"d-fulls\"><details class=\"d-full\"><summary>View full Codex output<\/summary><div class=\"d-full-pane\"><p class=\"d-full-note\">Complete sanitized first-valid output<\/p><pre>Implemented the scheduled-job priority feature.\n\nFiles changed:\n\n- [scheduler\/models.py](.\/scheduler\/models.py:5) \u2014 added serialized `priority`, defaulting to `0`.\n- [scheduler\/service.py](.\/scheduler\/service.py:11) \u2014 added validation, storage, and `list_jobs(min_priority=0)` filtering while preserving identifier order.\n- [tests\/test_scheduler.py](.\/tests\/test_scheduler.py:16) \u2014 added coverage for defaults, serialization, storage, boundaries, invalid values, and filtering.\n\nTests run:\n\n```text\npython3 -m unittest discover -s tests -v\nRan 7 tests \u2014 OK\n```\n\n**Conclusion:** The frozen task is complete and verified.  \n**Decision required:** None.<\/pre><\/div><\/details><details class=\"d-full\"><summary>View full Claude Code output<\/summary><div class=\"d-full-pane\"><p class=\"d-full-note\">Complete handoff record including the verbatim first-valid output<\/p><pre># T2 Run Record \u2014 Multi-File Priority Feature\n\n## Execution parameters\n\n| Field | Value |\n| --- | --- |\n| Task | `T2` |\n| Frozen prompt | `01-prompts\/T2.md` |\n| Frozen prompt SHA-256 | `8476a56034c4e719b97ee4081b32b4c73efd45feafe1bda9d477c83cb8fdb698` |\n| Execution unit | isolated Claude subagent, `general-purpose`, fresh context |\n| Isolation method | **subagent fallback** (documented downgrade \u2014 see `results\/isolation-decision.md`) |\n| Assigned clean copy | `runtime\/workspaces\/T2` |\n| Claude Code version | `2.1.220` |\n| Model setting | `claude-opus-5` |\n| start_utc | `2026-07-28T10:27:52Z` |\n| end_utc | `2026-07-28T10:29:33Z` |\n| Duration | 101 s (concurrent with T1, T3, T4; not a clean latency measurement) |\n| Attempt number | 1 |\n| Status | `valid` |\n| Human intervention | none |\n\n## Safe tool summary\n\n`date`, `Read`\/`Edit`\/`Write` inside the assigned workspace only, `grep`,\n`python3 -m unittest discover`, `python3 -m compileall`, inline `python3 -c` smoke checks.\nNo network, no installs, standard library only.\n\n## Files changed\n\n`scheduler\/models.py`, `scheduler\/repository.py`, `scheduler\/service.py`,\n`tests\/test_scheduler.py`. No files added or removed, so `fixture-manifest.json` remains\naccurate. Full patch: `results\/T2\/patch.diff`.\n\n## Offline check results\n\nReported by the task runner:\n\n| Check | Result |\n| --- | --- |\n| `python3 -m unittest discover -s tests -v` (baseline, pre-change) | 2 passed, 0 failed |\n| `python3 -m compileall -q scheduler tests` | clean |\n| inline smoke checks (legacy 3-arg construction, keyword\/positional threshold, `to_dict`) | passed |\n| `python3 -m unittest discover -s tests -v` (final) | **14 passed, 0 failed** |\n\nIndependently re-verified by the coordinator from outside the workspace:\n\n| Check | Exit | Result |\n| --- | --- | --- |\n| public suite (`results\/checks\/T2-public.txt`) | 0 | 14 tests, OK |\n| hidden validator `03-checks\/test_t2_hidden.py` (`results\/checks\/T2-hidden.txt`) | 0 | 2 tests, OK |\n\nThe hidden validator fails (exit 1) against a pristine fixture copy, confirming the pass is\nmeaningful. See `results\/checks\/validator-controls.md`.\n\n## Gold-map behavior coverage\n\nAll five required behaviors in `03-checks\/gold-map.json` are satisfied:\n\n| Required behavior | Evidence |\n| --- | --- |\n| priority default 0 | `Job.priority: int = MIN_PRIORITY`; `create_job(..., priority=MIN_PRIORITY)`; `test_priority_defaults_to_zero` |\n| integer range 0 through 9 | `validate_priority` rejects non-int and out-of-range; `test_out_of_range_priority_is_rejected`, `test_non_integer_priority_is_rejected` |\n| serialized priority | `to_dict()` via `asdict`; `test_priority_appears_in_serialized_output` asserts the exact dict |\n| minimum priority filtering | `list_jobs(min_priority=0)` \u2192 `JobRepository.all(min_priority)`; `test_filter_is_inclusive_of_the_threshold` |\n| identifier ordering | filter applied inside the existing `sorted(self._jobs)` comprehension; `test_filter_preserves_ascending_identifier_order` |\n\n## Sanitized errors\n\nNone. `__pycache__` artifacts from test runs were removed by the task runner; a coordinator\nsweep confirmed the workspace is clean.\n\n## Visible output (verbatim)\n\nAdded an integer `priority` field to scheduled jobs across the model, repository, and service layers of the local scheduler fixture, plus filtered listing.\n\n**`scheduler\/models.py`** \u2014 Added module constants `MIN_PRIORITY = 0` \/ `MAX_PRIORITY = 9` and a reusable `validate_priority(priority)` function that raises `ValueError` for non-integers and for values outside 0\u20139 inclusive. `bool` is explicitly rejected despite being an `int` subclass, so `True` cannot slip through as priority 1. The frozen `Job` dataclass gains `priority: int = MIN_PRIORITY` as a trailing field with a default, which keeps existing 3-argument construction working, and a `__post_init__` that calls the validator so invalid priorities cannot be constructed by any path. `to_dict()` was left as `asdict(self)`, so priority appears in serialized output automatically and in field order.\n\n**`scheduler\/repository.py`** \u2014 `JobRepository.all()` takes an optional `min_priority` and filters within the existing `sorted(self._jobs)` comprehension, so ascending identifier order is guaranteed by the same single code path that already provided it rather than being re-established after filtering. The default argument preserves the previous no-arg behavior for existing callers. Storage itself needed no change: the repository stores whole `Job` instances, so priority persists and round-trips through `save`\/`get`\/`all`.\n\n**`scheduler\/service.py`** \u2014 `create_job` accepts an optional trailing `priority` argument (default 0) and validates it before constructing the `Job`, so the failure is raised at the service boundary with the same `ValueError` contract as the existing blank-name check. `list_jobs(min_priority=0)` delegates to the repository. Ordering and filtering live in one place rather than being duplicated.\n\n**Design note on `min_priority` validation:** it is type-checked as an integer but deliberately *not* range-restricted to 0\u20139. The 0\u20139 rule constrains what a job may store; a threshold is a query parameter, and rejecting e.g. `min_priority=10` would turn a legitimate \"match nothing\" query into an error. `list_jobs(min_priority=9)` returning `[]` when no job has priority 9 is covered by a test.\n\n`scheduler\/time_rules.py` was read but left untouched \u2014 its documented daylight-saving shift is pre-existing seeded behavior, out of scope for this task, and changing it would have been an unrelated behavior change. No files were added or removed, so `fixture-manifest.json` remains accurate (verified programmatically).\n\n### Added test coverage\n\nDefault of 0; storage round-trip through `JobRepository`; priority in `to_dict()` output; boundary values 0 and 9 accepted; out-of-range \u22121\/10\/100 rejected; non-integer `\"3\"`\/`3.5`\/`None`\/`True` rejected; direct model-level validation; unfiltered listing order; filtered listing preserving ascending identifier order; inclusive threshold; empty result; non-integer `min_priority` rejected.\n<\/pre><\/div><\/details><\/div><p class=\"d-limit\">Fixture limitation: this was a small offline Python repository. Claude used a subagent fallback after CLI OAuth expired, so isolation and token visibility were not identical.<\/p><\/article>\n\n\n\n<h3 class=\"wp-block-heading\">Bug fixing under edge cases<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">T3 tested a daylight-saving-time scheduling defect. Both agents earned 30\/30 and passed the same spring and autumn hidden checks. Claude Code added explicit coverage for ambiguous and skipped local times, naive inputs, and multi-day rolls. Codex made a narrower calendar-day change and passed the frozen contract.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Broader edge-case tests create confidence when the library is widely reused or the cost of a scheduling error is high. A narrow change is attractive when the defect is well isolated and the team wants to minimize behavior changes. The practical review question is not \u201cWhich wrote more?\u201d It is whether the patch fixes the stated root cause, preserves surrounding contracts, and proves the risky boundaries.<\/p>\n\n\n\n<style>\n.evidence-t3-dst-bug-fix{--ink:#111318;--body:#414753;--muted:#707784;--line:#e4e7ec;--soft:#f6f7f9;--blue:#315efb;--violet:#7257e8;--green:#087a55;--greenSoft:#eaf8f2;--red:#a23131;--redSoft:#fff0f0;color:var(--ink);font-family:Inter,ui-sans-serif,-apple-system,BlinkMacSystemFont,\"Segoe UI\",sans-serif;line-height:1.5;border:1px solid var(--line);border-radius:8px;background:#fff;overflow:hidden;box-shadow:0 16px 42px rgba(24,32,56,.08)}\n.evidence-t3-dst-bug-fix *{box-sizing:border-box}.evidence-t3-dst-bug-fix .d-head{display:flex;align-items:flex-start;justify-content:space-between;gap:22px;padding:24px 26px 20px;border-bottom:1px solid var(--line)}\n.evidence-t3-dst-bug-fix .d-kicker{margin:0 0 6px;color:var(--muted);font:750 11px\/1.4 ui-monospace,SFMono-Regular,Menlo,monospace;letter-spacing:.08em;text-transform:uppercase}.evidence-t3-dst-bug-fix .d-title{margin:0;font-size:22px;line-height:1.2;letter-spacing:-.025em}.evidence-t3-dst-bug-fix .d-sub{margin:7px 0 0;color:var(--body);font-size:13px}.evidence-t3-dst-bug-fix .d-pill{flex:0 0 auto;padding:6px 10px;border:1px solid #ccd6ff;border-radius:999px;color:#2448ca;background:#f0f3ff;font-size:11px;font-weight:800}.evidence-t3-dst-bug-fix .d-grid{display:grid;grid-template-columns:1fr 1fr}.evidence-t3-dst-bug-fix .d-side{min-width:0;padding:24px 26px}.evidence-t3-dst-bug-fix .d-side+.d-side{border-left:1px solid var(--line)}.evidence-t3-dst-bug-fix .d-product{display:flex;align-items:baseline;justify-content:space-between;gap:14px;font-weight:800}.evidence-t3-dst-bug-fix .d-score{font-size:34px;line-height:1;letter-spacing:-.05em}.evidence-t3-dst-bug-fix .d-score small{color:var(--muted);font-size:12px;letter-spacing:0}.evidence-t3-dst-bug-fix .d-metrics{display:grid;grid-template-columns:repeat(3,1fr);gap:7px;margin:18px 0}.evidence-t3-dst-bug-fix .d-metric{padding:10px;border:1px solid var(--line);border-radius:6px;background:var(--soft)}.evidence-t3-dst-bug-fix .d-metric span{display:block;color:var(--muted);font-size:9px;font-weight:800;letter-spacing:.06em;text-transform:uppercase}.evidence-t3-dst-bug-fix .d-metric strong{display:block;margin-top:3px;font-size:13px}.evidence-t3-dst-bug-fix .d-preview{position:relative;height:192px;overflow:hidden;border:1px solid #29313f;border-radius:6px;background:#11151d;color:#edf0f5}.evidence-t3-dst-bug-fix .d-preview-label{padding:11px 13px;border-bottom:1px solid rgba(255,255,255,.1);color:#aeb7c7;font:750 9px\/1 ui-monospace,SFMono-Regular,Menlo,monospace;letter-spacing:.06em;text-transform:uppercase}.evidence-t3-dst-bug-fix .d-preview pre,.evidence-t3-dst-bug-fix .d-full pre{margin:0;white-space:pre-wrap;font:11px\/1.6 ui-monospace,SFMono-Regular,Menlo,Consolas,monospace}.evidence-t3-dst-bug-fix .d-preview pre{padding:13px}.evidence-t3-dst-bug-fix .d-preview:after{content:\"\";position:absolute;left:0;right:0;bottom:0;height:55px;background:linear-gradient(transparent,#11151d)}.evidence-t3-dst-bug-fix .d-take{margin:14px 0 0;color:var(--body);font-size:12px}.evidence-t3-dst-bug-fix .d-fulls{padding:0 26px 22px}.evidence-t3-dst-bug-fix .d-full{margin-top:10px;border:1px solid var(--line);border-radius:6px}.evidence-t3-dst-bug-fix .d-full summary{cursor:pointer;padding:13px 14px;font-size:12px;font-weight:800;list-style:none}.evidence-t3-dst-bug-fix .d-full summary::-webkit-details-marker{display:none}.evidence-t3-dst-bug-fix .d-full summary:after{content:\"+\";float:right;color:var(--blue);font-size:17px}.evidence-t3-dst-bug-fix .d-full[open] summary:after{content:\"\u2212\"}.evidence-t3-dst-bug-fix .d-full-pane{padding:0 14px 14px}.evidence-t3-dst-bug-fix .d-full-note{margin:0 0 9px;color:var(--muted);font-size:10px}.evidence-t3-dst-bug-fix .d-full pre{max-height:620px;overflow:auto;padding:15px;border-radius:5px;background:#f5f6f8;color:#222936;border:1px solid #e0e3e8}.evidence-t3-dst-bug-fix .d-limit{margin:0;padding:12px 26px;border-top:1px solid var(--line);color:#5a6270;background:#fafbfc;font-size:11px}\n@media(max-width:760px){.evidence-t3-dst-bug-fix .d-head{display:block}.evidence-t3-dst-bug-fix .d-pill{display:inline-block;margin-top:12px}.evidence-t3-dst-bug-fix .d-grid{grid-template-columns:1fr}.evidence-t3-dst-bug-fix .d-side+.d-side{border-left:0;border-top:1px solid var(--line)}}@media(max-width:480px){.evidence-t3-dst-bug-fix .d-head,.evidence-t3-dst-bug-fix .d-side{padding:20px}.evidence-t3-dst-bug-fix .d-fulls{padding:0 20px 20px}.evidence-t3-dst-bug-fix .d-metrics{grid-template-columns:1fr}}\n<\/style><article class=\"evidence-t3-dst-bug-fix\"><header class=\"d-head\"><div><p class=\"d-kicker\">T3 \u00b7 Controlled repository test<\/p><h3 class=\"d-title\">Repair daily scheduling across DST transitions<\/h3><p class=\"d-sub\">Preserve the requested New York wall-clock hour through spring and fall transitions.<\/p><\/div><span class=\"d-pill\">Tie \u00b7 different strengths<\/span><\/header><div class=\"d-grid\"><section class=\"d-side\" aria-label=\"Codex result\"><div class=\"d-product\"><span>Codex<\/span><span class=\"d-score\">30<small>\/30<\/small><\/span><\/div><div class=\"d-metrics\"><div class=\"d-metric\"><span>Duration<\/span><strong>99 s<\/strong><\/div><div class=\"d-metric\"><span>Checks<\/span><strong>4 + 2 hidden<\/strong><\/div><div class=\"d-metric\"><span>Scope<\/span><strong>Narrow fix<\/strong><\/div><\/div><div class=\"d-preview\"><div class=\"d-preview-label\">First-valid output excerpt<\/div><pre>Root cause: the implementation added 24 elapsed UTC hours.\n\nFix: advance by one local calendar day. Both frozen spring and fall checks passed.<\/pre><\/div><p class=\"d-take\"><strong>Reviewer take:<\/strong> The narrower patch passed the complete frozen contract.<\/p><\/section><section class=\"d-side\" aria-label=\"Claude Code result\"><div class=\"d-product\"><span>Claude Code<\/span><span class=\"d-score\">30<small>\/30<\/small><\/span><\/div><div class=\"d-metrics\"><div class=\"d-metric\"><span>Duration<\/span><strong>118 s<\/strong><\/div><div class=\"d-metric\"><span>Checks<\/span><strong>10 + 2 hidden<\/strong><\/div><div class=\"d-metric\"><span>Scope<\/span><strong>Edge coverage<\/strong><\/div><\/div><div class=\"d-preview\"><div class=\"d-preview-label\">First-valid output excerpt<\/div><pre>Covered spring, fall, skipped and ambiguous local times, naive input, and multi-day rolls.\n\nBoth frozen hidden checks passed after broader regression coverage.<\/pre><\/div><p class=\"d-take\"><strong>Reviewer take:<\/strong> Broader edge-case coverage; the same functional score under the frozen rubric.<\/p><\/section><\/div><div class=\"d-fulls\"><details class=\"d-full\"><summary>View full Codex output<\/summary><div class=\"d-full-pane\"><p class=\"d-full-note\">Complete sanitized first-valid output<\/p><pre>Implemented the DST fix.\n\nFiles changed:\n\n- [scheduler\/time_rules.py](.\/scheduler\/time_rules.py)\n- [tests\/test_time_rules.py](.\/tests\/test_time_rules.py)\n\nChanges:\n\n- Advances one local calendar day instead of 24 UTC hours.\n- Added New York spring-forward and fall-back regression tests.\n- Verified elapsed UTC intervals are 23 and 25 hours respectively.\n\nChecks:\n\n- Unit tests: 4 passed.\n- In-memory syntax check: 7 Python files passed.\n- Bytecode compilation was blocked because Python redirects its cache outside the assigned workspace; no external files were written.\n\n**Conclusion:** `next_daily_run` now preserves the requested wall-clock hour across DST transitions. No decision is required.<\/pre><\/div><\/details><details class=\"d-full\"><summary>View full Claude Code output<\/summary><div class=\"d-full-pane\"><p class=\"d-full-note\">Complete handoff record including the verbatim first-valid output<\/p><pre># T3 Run Record \u2014 DST Bug Fix\n\n## Execution parameters\n\n| Field | Value |\n| --- | --- |\n| Task | `T3` |\n| Frozen prompt | `01-prompts\/T3.md` |\n| Frozen prompt SHA-256 | `3fee50d0230f27f8d6dd7f12e050fa3f0b7f5fb77d9559361fee96e5f1a47307` |\n| Execution unit | isolated Claude subagent, `general-purpose`, fresh context |\n| Isolation method | **subagent fallback** (documented downgrade \u2014 see `results\/isolation-decision.md`) |\n| Assigned clean copy | `runtime\/workspaces\/T3` |\n| Claude Code version | `2.1.220` |\n| Model setting | `claude-opus-5` |\n| start_utc | `2026-07-28T10:28:05Z` |\n| end_utc | `2026-07-28T10:30:03Z` |\n| Duration | 118 s (concurrent with T1, T2, T4; not a clean latency measurement) |\n| Attempt number | 1 |\n| Status | `valid` |\n| Human intervention | none |\n\n## Safe tool summary\n\n`date`, `Read`\/`Edit` inside the assigned workspace only, `python3 -m unittest discover`,\n`python3 -m compileall`, an in-memory negative-control run, a manifest integrity check.\nStandard library only (`datetime`, `zoneinfo`, `unittest`). No network, no installs.\n\n## Files changed\n\n`scheduler\/time_rules.py` (the fix \u2014 only the body and docstring of `next_daily_run`),\n`tests\/test_scheduler.py` (added `NextDailyRunTests`, 8 tests, plus an `_elapsed` helper;\nexisting tests unmodified). No other module touched, no files added. Full patch:\n`results\/T3\/patch.diff`.\n\n## Offline check results\n\nReported by the task runner:\n\n| Check | Result |\n| --- | --- |\n| `python3 -m unittest discover -s tests -v` | **10 passed, 0 failed** (2 pre-existing + 8 new) |\n| `python3 -m compileall -q scheduler tests` | clean |\n| manifest integrity vs `fixture-manifest.json` | 0 missing, 0 unexpected |\n| negative control: seeded implementation patched in-memory, `NextDailyRunTests` only | 8 ran, **5 failed as intended**, including both DST tests |\n\nIndependently re-verified by the coordinator from outside the workspace:\n\n| Check | Exit | Result |\n| --- | --- | --- |\n| public suite (`results\/checks\/T3-public.txt`) | 0 | 10 tests, OK |\n| hidden validator `03-checks\/test_t3_hidden.py` (`results\/checks\/T3-hidden.txt`) | 0 | 2 tests, OK |\n\nThe hidden validator fails (exit 1) against a pristine fixture copy. See\n`results\/checks\/validator-controls.md`.\n\n## Gold-map behavior coverage\n\nAll four required behaviors in `03-checks\/gold-map.json` are satisfied:\n\n| Required behavior | Evidence |\n| --- | --- |\n| wall-clock hour preserved | increment applied to the naive local wall clock, zone re-attached; asserted in every DST test |\n| spring transition covered | `test_spring_forward_keeps_requested_hour` \u2014 hour 09:30 held, offset `-04:00`, elapsed 23 h |\n| fall transition covered | `test_fall_back_keeps_requested_hour` \u2014 hour 09:30 held, offset `-05:00`, elapsed 25 h |\n| IANA timezone used | `ZoneInfo(\"America\/New_York\")` from the standard library |\n\nBeyond the required set, the runner also covered the two PEP 495 edge cases \u2014 a wall time\nskipped by spring-forward and an ambiguous wall time during fall-back \u2014 plus seven-day rolls\nstraddling each transition and naive-input behavior.\n\n## Sanitized errors\n\nTwo self-inflicted test-assertion failures: `following - previous` on two datetimes sharing a\n`tzinfo` subtracts naive wall clocks, so the elapsed-time assertions read 1 day and failed.\nThe implementation was correct and the assertions were wrong; corrected with an `_elapsed`\nhelper that converts both operands to UTC first. No tooling, import, or environment errors.\n`__pycache__` artifacts were removed after the run.\n\n## Visible output (verbatim)\n\n**Reproduction**\n\n`scheduler\/time_rules.py::next_daily_run` was called with an `America\/New_York` datetime on each side of a 2024 DST transition:\n\n```\n2024-03-09T09:30:00-05:00 -&gt; 2024-03-10T10:30:00-04:00   # hour drifted +1\n2024-11-02T09:30:00-04:00 -&gt; 2024-11-03T08:30:00-05:00   # hour drifted -1\n```\n\nA job requested for 09:30 local fired at 10:30 after spring-forward and 08:30 after fall-back.\n\n**Root cause**\n\nThe seeded implementation treated \"daily\" as a fixed absolute duration:\n\n```python\nelapsed = previous_run.astimezone(timezone.utc) + timedelta(hours=24)\nreturn elapsed.astimezone(previous_run.tzinfo)\n```\n\nConverting to UTC first makes `timedelta(hours=24)` exactly 24 elapsed hours. But a local calendar day in an IANA zone is 23 h on the spring-forward day and 25 h on the fall-back day, so a constant 24 h offset necessarily lands one hour away from the requested wall-clock time.\n\n**The fix** (`scheduler\/time_rules.py`)\n\nDaily recurrence is a calendar rule, so the increment is now applied to the naive local wall clock, with the zone re-attached afterwards:\n\n```python\ntzinfo = previous_run.tzinfo\nnaive_next = previous_run.replace(tzinfo=None) + timedelta(days=1)\nif tzinfo is None:\n    return naive_next\ncandidate = naive_next.replace(tzinfo=tzinfo, fold=0)\nreturn candidate.astimezone(timezone.utc).astimezone(tzinfo)\n```\n\nWhy it is correct:\n- Wall-clock arithmetic on the naive value is offset-independent, so the requested hour:minute is preserved by construction; `zoneinfo` then supplies whichever UTC offset is in force on the new date (verified: `-04:00` after spring-forward, `-05:00` after fall-back, elapsed real time 23 h and 25 h respectively).\n- `fold=0` plus the UTC round-trip handles the two edge cases that pure wall-clock arithmetic leaves ill-defined. Under PEP 495, a *skipped* wall time (02:30 on 2024-03-10) has the pre-transition offset at `fold=0`, so the round-trip normalizes it to the real instant 03:30 EDT rather than returning a datetime that names no instant. An *ambiguous* wall time (01:30 on 2024-11-03) resolves deterministically to its first occurrence (EDT) instead of being arbitrary.\n- Naive input keeps its previous meaning (plain +1 calendar day) \u2014 no new `tzinfo` requirement is imposed on callers.\n\nScope was kept narrow: only the body\/docstring of `next_daily_run` changed. `service.py`, `models.py`, `repository.py`, and `__init__.py` are untouched, no new files were added, and the fixture manifest still matches exactly (no missing or unexpected files). Standard library only (`datetime`, `zoneinfo`, `unittest`).\n\n**Regression tests** (`tests\/test_scheduler.py`, new `NextDailyRunTests`, 8 tests) cover: ordinary-day stability, spring-forward hour + offset + 23 h elapsed, fall-back hour + offset + 25 h elapsed, seven-day rolls straddling each transition, skipped wall time, ambiguous wall time, and naive input.\n\nTwo notes worth recording. First, my initial elapsed-time assertions used `following - previous`; Python ignores a shared `tzinfo` and subtracts naive wall clocks, so both returned 1 day and failed. The implementation was right and the assertions wrong \u2014 fixed with an `_elapsed` helper that converts both operands to UTC. Second, I verified the tests actually catch the original defect by re-running `NextDailyRunTests` against an in-memory copy of the seeded implementation: 5 of 8 failed, including both DST tests. The three that pass under the buggy version (ordinary day, skipped, ambiguous) are behavior-documenting rather than regression-catching \u2014 the skipped\/ambiguous cases coincidentally agree because a +24 h absolute shift happens to land on the same normalized instant.\n<\/pre><\/div><\/details><\/div><p class=\"d-limit\">Fixture limitation: this was a small offline Python repository. Claude used a subagent fallback after CLI OAuth expired, so isolation and token visibility were not identical.<\/p><\/article>\n\n\n\n<h2 id=\"code-review\" class=\"wp-block-heading\">Code review quality and human verification<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Code-review tasks expose a different skill from feature implementation. A useful reviewer must find reproducible issues, rank their impact, cite the exact code path, and distinguish a real defect from a speculative warning. Fixing an issue adds another test: the patch must not introduce a new regression.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Claude Code had the clearest advantage in T4. It reported five independently reproducible findings, including both seeded review targets, and earned 19\/20. Codex reported and fixed two valid higher-severity defects but did not mention the seeded pair, earning 18\/20. Codex&#8217;s findings were not false positives; the observed difference was breadth rather than basic correctness.<\/p>\n\n\n\n<style>\n.evidence-t4-review-and-repair{--ink:#111318;--body:#414753;--muted:#707784;--line:#e4e7ec;--soft:#f6f7f9;--blue:#315efb;--violet:#7257e8;--green:#087a55;--greenSoft:#eaf8f2;--red:#a23131;--redSoft:#fff0f0;color:var(--ink);font-family:Inter,ui-sans-serif,-apple-system,BlinkMacSystemFont,\"Segoe UI\",sans-serif;line-height:1.5;border:1px solid var(--line);border-radius:8px;background:#fff;overflow:hidden;box-shadow:0 16px 42px rgba(24,32,56,.08)}\n.evidence-t4-review-and-repair *{box-sizing:border-box}.evidence-t4-review-and-repair .d-head{display:flex;align-items:flex-start;justify-content:space-between;gap:22px;padding:24px 26px 20px;border-bottom:1px solid var(--line)}\n.evidence-t4-review-and-repair .d-kicker{margin:0 0 6px;color:var(--muted);font:750 11px\/1.4 ui-monospace,SFMono-Regular,Menlo,monospace;letter-spacing:.08em;text-transform:uppercase}.evidence-t4-review-and-repair .d-title{margin:0;font-size:22px;line-height:1.2;letter-spacing:-.025em}.evidence-t4-review-and-repair .d-sub{margin:7px 0 0;color:var(--body);font-size:13px}.evidence-t4-review-and-repair .d-pill{flex:0 0 auto;padding:6px 10px;border:1px solid #ccd6ff;border-radius:999px;color:#2448ca;background:#f0f3ff;font-size:11px;font-weight:800}.evidence-t4-review-and-repair .d-grid{display:grid;grid-template-columns:1fr 1fr}.evidence-t4-review-and-repair .d-side{min-width:0;padding:24px 26px}.evidence-t4-review-and-repair .d-side+.d-side{border-left:1px solid var(--line)}.evidence-t4-review-and-repair .d-product{display:flex;align-items:baseline;justify-content:space-between;gap:14px;font-weight:800}.evidence-t4-review-and-repair .d-score{font-size:34px;line-height:1;letter-spacing:-.05em}.evidence-t4-review-and-repair .d-score small{color:var(--muted);font-size:12px;letter-spacing:0}.evidence-t4-review-and-repair .d-metrics{display:grid;grid-template-columns:repeat(3,1fr);gap:7px;margin:18px 0}.evidence-t4-review-and-repair .d-metric{padding:10px;border:1px solid var(--line);border-radius:6px;background:var(--soft)}.evidence-t4-review-and-repair .d-metric span{display:block;color:var(--muted);font-size:9px;font-weight:800;letter-spacing:.06em;text-transform:uppercase}.evidence-t4-review-and-repair .d-metric strong{display:block;margin-top:3px;font-size:13px}.evidence-t4-review-and-repair .d-preview{position:relative;height:192px;overflow:hidden;border:1px solid #29313f;border-radius:6px;background:#11151d;color:#edf0f5}.evidence-t4-review-and-repair .d-preview-label{padding:11px 13px;border-bottom:1px solid rgba(255,255,255,.1);color:#aeb7c7;font:750 9px\/1 ui-monospace,SFMono-Regular,Menlo,monospace;letter-spacing:.06em;text-transform:uppercase}.evidence-t4-review-and-repair .d-preview pre,.evidence-t4-review-and-repair .d-full pre{margin:0;white-space:pre-wrap;font:11px\/1.6 ui-monospace,SFMono-Regular,Menlo,Consolas,monospace}.evidence-t4-review-and-repair .d-preview pre{padding:13px}.evidence-t4-review-and-repair .d-preview:after{content:\"\";position:absolute;left:0;right:0;bottom:0;height:55px;background:linear-gradient(transparent,#11151d)}.evidence-t4-review-and-repair .d-take{margin:14px 0 0;color:var(--body);font-size:12px}.evidence-t4-review-and-repair .d-fulls{padding:0 26px 22px}.evidence-t4-review-and-repair .d-full{margin-top:10px;border:1px solid var(--line);border-radius:6px}.evidence-t4-review-and-repair .d-full summary{cursor:pointer;padding:13px 14px;font-size:12px;font-weight:800;list-style:none}.evidence-t4-review-and-repair .d-full summary::-webkit-details-marker{display:none}.evidence-t4-review-and-repair .d-full summary:after{content:\"+\";float:right;color:var(--blue);font-size:17px}.evidence-t4-review-and-repair .d-full[open] summary:after{content:\"\u2212\"}.evidence-t4-review-and-repair .d-full-pane{padding:0 14px 14px}.evidence-t4-review-and-repair .d-full-note{margin:0 0 9px;color:var(--muted);font-size:10px}.evidence-t4-review-and-repair .d-full pre{max-height:620px;overflow:auto;padding:15px;border-radius:5px;background:#f5f6f8;color:#222936;border:1px solid #e0e3e8}.evidence-t4-review-and-repair .d-limit{margin:0;padding:12px 26px;border-top:1px solid var(--line);color:#5a6270;background:#fafbfc;font-size:11px}\n@media(max-width:760px){.evidence-t4-review-and-repair .d-head{display:block}.evidence-t4-review-and-repair .d-pill{display:inline-block;margin-top:12px}.evidence-t4-review-and-repair .d-grid{grid-template-columns:1fr}.evidence-t4-review-and-repair .d-side+.d-side{border-left:0;border-top:1px solid var(--line)}}@media(max-width:480px){.evidence-t4-review-and-repair .d-head,.evidence-t4-review-and-repair .d-side{padding:20px}.evidence-t4-review-and-repair .d-fulls{padding:0 20px 20px}.evidence-t4-review-and-repair .d-metrics{grid-template-columns:1fr}}\n<\/style><article class=\"evidence-t4-review-and-repair\"><header class=\"d-head\"><div><p class=\"d-kicker\">T4 \u00b7 Controlled repository test<\/p><h3 class=\"d-title\">Find reproducible defects and repair the highest-confidence issues<\/h3><p class=\"d-sub\">Review the fixture, reproduce real defects, apply permitted fixes, and prove regression tests bind the behavior.<\/p><\/div><span class=\"d-pill\">Claude coverage edge<\/span><\/header><div class=\"d-grid\"><section class=\"d-side\" aria-label=\"Codex result\"><div class=\"d-product\"><span>Codex<\/span><span class=\"d-score\">18<small>\/20<\/small><\/span><\/div><div class=\"d-metrics\"><div class=\"d-metric\"><span>Duration<\/span><strong>223 s<\/strong><\/div><div class=\"d-metric\"><span>Checks<\/span><strong>4 passed<\/strong><\/div><div class=\"d-metric\"><span>Scope<\/span><strong>2 fixed<\/strong><\/div><\/div><div class=\"d-preview\"><div class=\"d-preview-label\">First-valid output excerpt<\/div><pre>1. High \u2014 Existing jobs could be silently overwritten.\n\n2. Medium \u2014 Daily runs drifted across daylight-saving transitions.\n\nVerification: four tests passed, zero failures.<\/pre><\/div><p class=\"d-take\"><strong>Reviewer take:<\/strong> Useful higher-severity findings and focused fixes, but it missed the two seeded review targets.<\/p><\/section><section class=\"d-side\" aria-label=\"Claude Code result\"><div class=\"d-product\"><span>Claude Code<\/span><span class=\"d-score\">19<small>\/20<\/small><\/span><\/div><div class=\"d-metrics\"><div class=\"d-metric\"><span>Duration<\/span><strong>235 s<\/strong><\/div><div class=\"d-metric\"><span>Checks<\/span><strong>9 passed<\/strong><\/div><div class=\"d-metric\"><span>Scope<\/span><strong>5 reported<\/strong><\/div><\/div><div class=\"d-preview\"><div class=\"d-preview-label\">First-valid output excerpt<\/div><pre>D1 DST drift \u00b7 High\nD2 Naive timezone acquisition \u00b7 Medium\nD3 Identifier collision \u00b7 High\nD4 Missing delete raises KeyError \u00b7 Medium\nD5 Owner unvalidated \u00b7 Low<\/pre><\/div><p class=\"d-take\"><strong>Reviewer take:<\/strong> Broader review coverage with zero reproduced false positives; two contract-dependent findings were not fixed.<\/p><\/section><\/div><div class=\"d-fulls\"><details class=\"d-full\"><summary>View full Codex output<\/summary><div class=\"d-full-pane\"><p class=\"d-full-note\">Complete sanitized first-valid output<\/p><pre>## Actionable findings\n\n1. **High \u2014 Existing jobs could be silently overwritten**\n\n   - Evidence: identifier initialization in [service.py](.\/scheduler\/service.py:8), combined with dictionary replacement in [repository.py](.\/scheduler\/repository.py:9).\n   - Impact: creating a service with an existing repository could reuse identifier `1`, replacing the stored job and causing data loss.\n   - Minimal reproduction: save `Job(1, ...)`, construct `SchedulerService(repository)`, then create another job. Before the fix, only the new job remained.\n   - Fix: initialize the next identifier from the repository\u2019s highest existing identifier.\n   - Regression test: [test_scheduler.py](.\/tests\/test_scheduler.py:22).\n\n2. **Medium \u2014 Daily runs drifted across daylight-saving transitions**\n\n   - Evidence: corrected calculation in [time_rules.py](.\/scheduler\/time_rules.py:4).\n   - Impact: a daily job scheduled for 09:00 ran at 10:00 after the spring DST transition.\n   - Minimal reproduction: calculate the next run after `2026-03-07 09:00 America\/New_York`. Before the fix, the result was `2026-03-08 10:00`.\n   - Fix: add one local calendar day instead of 24 elapsed UTC hours.\n   - Regression test: [test_scheduler.py](.\/tests\/test_scheduler.py:33).\n\n## Files changed\n\n- [scheduler\/service.py](.\/scheduler\/service.py)\n- [scheduler\/time_rules.py](.\/scheduler\/time_rules.py)\n- [tests\/test_scheduler.py](.\/tests\/test_scheduler.py)\n\n## Verification\n\n- Both regression tests were observed failing before their fixes.\n- Each passed after its corresponding fix.\n- `PYTHONDONTWRITEBYTECODE=1 python3 -m unittest discover -s tests -v`\n- Result: **4 tests passed, 0 failures**.\n- No generated Python cache artifacts remain.\n\n## Conclusion and decision\n\n- **Conclusion:** Both confirmed defects are fixed and covered by focused regression tests.\n- **Decision required:** None.<\/pre><\/div><\/details><details class=\"d-full\"><summary>View full Claude Code output<\/summary><div class=\"d-full-pane\"><p class=\"d-full-note\">Complete handoff record including the verbatim first-valid output<\/p><pre># T4 Run Record \u2014 Test and Code Review\n\n## Execution parameters\n\n| Field | Value |\n| --- | --- |\n| Task | `T4` |\n| Frozen prompt | `01-prompts\/T4.md` |\n| Frozen prompt SHA-256 | `5e69546007ea776f40b3d334ddb75671ff684c151c77551bad27de3b0fdce640` |\n| Execution unit | isolated Claude subagent, `general-purpose`, fresh context |\n| Isolation method | **subagent fallback** (documented downgrade \u2014 see `results\/isolation-decision.md`) |\n| Assigned clean copy | `runtime\/workspaces\/T4` |\n| Claude Code version | `2.1.220` |\n| Model setting | `claude-opus-5` |\n| start_utc | `2026-07-28T10:28:22Z` |\n| end_utc | `2026-07-28T10:32:17Z` |\n| Duration | 235 s (concurrent with T1\u2013T3; not a clean latency measurement) |\n| Attempt number | 1 |\n| Status | `valid` |\n| Human intervention | none |\n\nThe conditional in the frozen prompt \u2014 \"fix them if the frozen run configuration permits\nfixes\" \u2014 was resolved by the coordinator from `00-control\/run-config.json` and\n`00-control\/safety-policy.json`, which permit writing isolated fixture copies. The runner was\ntold fixes are permitted. It was not shown any control file.\n\n## Safe tool summary\n\n`date`, `Read`\/`Edit`\/`Write` inside the assigned workspace only, `grep`,\n`python3 -m unittest discover`, `python3 -m compileall`, a scratch-copy vacuity run, a\nmanifest presence check. No network, no installs, standard library only.\n\n## Files changed\n\n`scheduler\/time_rules.py` (modified), `scheduler\/service.py` (modified),\n`tests\/test_scheduler.py` (modified, +2 tests), `tests\/test_time_rules.py` (added, 5 tests).\nFull patch: `results\/T4\/patch.diff`.\n\n## Offline check results\n\nReported by the task runner:\n\n| Check | Result |\n| --- | --- |\n| `python3 -m unittest discover -s tests -v` | **9 passed, 0 failed, 0 skipped** (baseline 2) |\n| `python3 -m compileall -q scheduler tests` | pass |\n| manifest presence check vs `expected_files` | all 8 present |\n| vacuity run against restored original implementations | 5 failures, 4 passed, as intended |\n\nIndependently re-verified by the coordinator from outside the workspace:\n\n| Check | Exit | Result |\n| --- | --- | --- |\n| public suite (`results\/checks\/T4-public.txt`) | 0 | 9 tests, OK |\n\nInformational cross-check, not scored: the T3 hidden validator also passes against this\nworkspace (exit 0), because T4 independently fixed the same `next_daily_run` defect. See\n`results\/checks\/validator-controls.md`.\n\n## Gold-map defect coverage\n\n`03-checks\/gold-map.json` seeds two defects for T4 scoring. Both were found:\n\n| Seeded defect | Found | Reported as |\n| --- | --- | --- |\n| blank owner accepted | yes | **D5**, `scheduler\/service.py:11,13`, severity Low, reported and deliberately not fixed |\n| delete of missing identifier raises `KeyError` | yes | **D4**, `scheduler\/repository.py:19` \u2192 `service.py:22`, severity Medium, reported and deliberately not fixed |\n\nThree additional findings were reported beyond the seeded set. None is a false positive:\n\n- **D1** \u2014 DST drift in `next_daily_run`. This is the third defect listed in\n  `03-checks\/seeded-defects.md` (it is scored under T3, not T4, but it is a genuine seeded bug).\n- **D2** \u2014 naive input to `next_daily_run` silently acquires the host timezone and returns an\n  aware datetime. Real and host-dependent; verified by the coordinator against the fixture source.\n- **D3** \u2014 identifier collision between a service and an injected or shared repository, causing\n  silent overwrite. Real: `SchedulerService._next_identifier` always starts at 1 while\n  `JobRepository.save` is an unguarded dict assignment. T1 independently identified the same\n  seam from the pristine fixture.\n\nFalse-positive count: **0**. The runner also correctly reported no security findings, noting\nthe fixture has no network, I\/O, deserialization, subprocess, or credential surface.\n\n## Deviation worth recording\n\nThe frozen prompt asked for tests and fixes on \"the two highest-confidence defects\". The\nrunner selected D1\/D2 (one fix, same function) and D3, rather than the two gold-map seeded\ndefects D4 and D5. Its stated reason is defensible: D4 and D5 both hinge on an unspecified API\ncontract, so fixing them would be a guess. Detection credit is unaffected \u2014 both seeded defects\nwere found, with file, line, severity, impact, and reproduction. The fixes applied were to real\ndefects and are backed by a vacuity check.\n\n## Sanitized errors\n\nTwo self-inflicted test failures from wall-clock versus absolute-time subtraction on datetimes\nsharing a `tzinfo`; corrected and re-verified. No tooling or environment errors. `__pycache__`\nartifacts were removed after the run.\n\n## Visible output (verbatim)\n\n### Defect report\n\n**D1 \u2014 `scheduler\/time_rules.py:10-11` \u2014 Daily schedule drifts by one hour across DST \u2014 Severity: High**\n\n`next_daily_run` adds 24 *elapsed* hours via a UTC round-trip instead of advancing the calendar day. A daily job is defined by wall-clock time, so every job in a DST-observing zone silently fires an hour early or late from the transition day onward, and the error persists (each run re-seeds from the shifted time). The docstring itself admitted the behavior.\n\n```\nprev = datetime(2024, 3, 9, 9, 0, tzinfo=ZoneInfo(\"America\/New_York\"))\nnext_daily_run(prev)   # -&gt; 2024-03-10 10:00:00-04:00, expected 09:00-04:00\nprev = datetime(2024, 11, 2, 9, 0, tzinfo=ZoneInfo(\"America\/New_York\"))\nnext_daily_run(prev)   # -&gt; 2024-11-03 08:00:00-05:00, expected 09:00-05:00\n```\n\n**D2 \u2014 `scheduler\/time_rules.py:10-11` \u2014 Naive input silently acquires the machine's local timezone \u2014 Severity: Medium**\n\nSame two lines. `.astimezone(timezone.utc)` on a naive datetime assumes system local time, and `.astimezone(None)` returns an aware datetime. A naive-in call returns aware-out, so the result is non-comparable with the caller's other naive values (`TypeError` on comparison) and the answer depends on the host's `TZ`.\n\n```\nnext_daily_run(datetime(2024, 1, 1, 9, 0))\n# -&gt; datetime(2024, 1, 2, 9, 0, tzinfo=timezone(timedelta(hours=8), 'CST'))  # host-dependent\n```\n\n**D3 \u2014 `scheduler\/service.py:8,13,15` (pre-fix) \u2014 Identifier collision silently destroys stored jobs \u2014 Severity: High**\n\nThe id counter was per-`SchedulerService` and always started at 1, ignoring the injected repository's contents. Any second service over a shared repository, or any service over a pre-populated repository, re-issued id 1; `JobRepository.save` is a dict assignment, so the existing job was overwritten with no error. Data loss.\n\n```\nrepo = JobRepository()\nSchedulerService(repo).create_job(\"Nightly backup\", \"Alex\")    # id 1\nSchedulerService(repo).create_job(\"Payroll export\", \"Blake\")   # id 1 again\nrepo.all()   # -&gt; [Job(1, 'Payroll export', 'Blake')]  \u2014 Alex's job is gone\n```\n\n**D4 \u2014 `scheduler\/repository.py:19` \u2192 `scheduler\/service.py:22` \u2014 `delete_job` raises `KeyError` for an unknown id \u2014 Severity: Medium**\n\n`delete` uses `dict.pop` without a default, while the sibling `get` (line 13) returns `None` for a miss. The inconsistency leaks a raw `KeyError` through `SchedulerService.delete_job`, which callers are unlikely to catch. Repro: `SchedulerService().delete_job(999)` \u2192 `KeyError: 999`. Fix would be `pop(identifier, None)` or an explicit domain exception \u2014 deliberately left unfixed, since the correct contract (silent no-op vs. typed error) is an API decision for the owner.\n\n**D5 \u2014 `scheduler\/service.py:11,13` \u2014 Owner field is unvalidated \u2014 Severity: Low**\n\n`name` is checked for blankness but `owner` is not: `create_job(\"ok\", \"   \")` stores `owner=''`. Both parameters also call `.strip()` unguarded, so `None` surfaces as `AttributeError: 'NoneType' object has no attribute 'strip'` rather than the `ValueError` the API otherwise uses for bad input. Left unfixed \u2014 the intended owner contract isn't stated anywhere in the fixture.\n\n**Security:** no findings. The fixture has no network, I\/O, deserialization, subprocess, or credential surface; storage is an in-process dict.\n\n**Coverage gap:** `scheduler\/time_rules.py` and `scheduler\/repository.py` had zero test coverage before this run; the two seeded tests exercised only the service happy path.\n\n### Tests and fixes applied\n\nThe two highest-confidence defects were **D1\/D2** (same function, one fix) and **D3**.\n\n**Fix 1 \u2014 `scheduler\/time_rules.py`:** replaced the UTC round-trip with `previous_run + timedelta(days=1)`. Python performs wall-clock arithmetic on aware datetimes, and `zoneinfo` re-resolves the UTC offset lazily from the new wall time \u2014 so 09:00 stays 09:00 while the offset correctly moves `-05:00 \u2192 -04:00`. This fixes D1 and D2 together: naive inputs now stay naive and the local zone is never inferred. The unreachable-vs-ambiguous transition-hour case is documented as deferring to the tzinfo's fold rules.\n\n**Fix 2 \u2014 `scheduler\/service.py`:** removed the per-instance `_next_identifier` counter; identifiers are now allocated from repository state as `max(existing ids, default=0) + 1` at create time. This is correct for both shared and pre-populated repositories. Trade-off: O(n) per create instead of O(1), acceptable for a dict-backed in-memory store, and correctness over a counter that cannot see the store it writes to.\n\n**Tests \u2014 `tests\/test_time_rules.py` (new, 5 tests):** spring-forward and fall-back wall-clock preservation, absolute elapsed time of 23h\/25h across those transitions, UTC-offset re-resolution, naive-stays-naive, and a plain UTC day advance. Skips cleanly if no IANA database is present.\n\n**Tests \u2014 `tests\/test_scheduler.py` (extended, +2 tests):** two services sharing one repository must issue distinct ids and retain both jobs; a service over a repository pre-populated with ids 1 and 7 must issue 8 and leave the existing jobs intact.\n\n**Vacuity check:** I restored both original implementations in a scratch copy and re-ran the suite \u2014 all 5 new tests failed (`FAILED (failures=5)`), and all pass against the fixed code. The tests genuinely bind the defects.\n\nOne correction during the run: my first draft asserted elapsed DST time as `result - previous`, which Python evaluates as wall-clock arithmetic when both operands share a tzinfo, so it read `24h` and the tests failed. The assertions now convert to UTC before subtracting. The failure was in my test, not in the fix.\n\n### Notes for the owner\n\n`fixture-manifest.json` was left untouched, so `tests\/test_time_rules.py` is a legitimate new file not listed in `expected_files` \u2014 update the manifest if a check enforces an exact file set. And D4\/D5 were reported but deliberately not fixed: both hinge on an API contract decision (miss-on-delete semantics, owner validation rules) that isn't specified anywhere in the fixture, so picking one would have been a guess rather than a fix.\n<\/pre><\/div><\/details><\/div><p class=\"d-limit\">Fixture limitation: this was a small offline Python repository. Claude used a subagent fallback after CLI OAuth expired, so isolation and token visibility were not identical.<\/p><\/article>\n\n\n\n<p class=\"wp-block-paragraph\">Neither score removes the need for human review. After either agent completes a repair, inspect the changed interfaces, unexpected file scope, missing negative tests, permission or command history, and rollback path. A confident explanation is not a substitute for a reproducible test.<\/p>\n\n\n\n<h2 id=\"safety\" class=\"wp-block-heading\">Permissions, sandboxing, and operational safety<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">A coding agent can read private code, execute shell commands, modify many files, and call external tools. That makes permission design part of product quality. The relevant questions are which operations require approval, what the agent can access by default, how clearly it previews risky actions, and whether a developer can inspect the resulting changes.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Do not treat fewer approval prompts as automatically better. Low friction is useful in an isolated fixture with no credentials or production connections. The same behavior can be dangerous in a repository connected to deployment scripts, real customer data, or broad cloud permissions. Conversely, excessive approvals can make safe automation impractical and encourage users to approve requests without reading them.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Our test used fresh local copies, no production systems, no deployment, and no valid-run human intervention. It therefore evaluates task completion under safe conditions, not the relative security of the two products. Before using either tool on important work, start with read-only access, define allowed directories and commands, review the diff, and run objective checks in a recoverable environment.<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Begin with a read-only architecture or risk pass.<\/li>\n\n\n\n<li>Approve only commands whose scope and consequence you understand.<\/li>\n\n\n\n<li>Keep credentials, production data, and deployment access outside the trial.<\/li>\n\n\n\n<li>Require a diff, tests, and a clear rollback path before accepting the result.<\/li>\n<\/ul>\n\n\n\n<h2 id=\"customization\" class=\"wp-block-heading\">MCP, Skills, project instructions, and customization<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Both ecosystems can be extended, but extension terms should not be treated as synonyms. Project instructions tell an agent how to behave in a repository. Reusable Skills package a repeatable workflow or expertise. MCP connects a host to external tools or data. Hooks and commands can automate particular moments in a development loop. A product can be strong in one layer without matching another feature one-for-one.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The buying question is not whether an integration logo exists. It is whether the connection can be installed, authenticated, approved, and exercised in the host you actually use. It also needs an understandable failure mode. An MCP server that works interactively but waits for approval in an unattended session is still useful, but it should not be described as frictionless automation.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Reusable instructions have a similar caveat. Long instruction files do not guarantee compliance, and an overly complicated rule set can consume context or create contradictions. Keep repository guidance short, testable, and close to the code. State required verification commands, protected areas, style rules, and when the agent must stop to ask.<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>Project instructions:<\/strong> repository-specific rules and verification commands.<\/li>\n\n\n\n<li><strong>Skills:<\/strong> reusable procedures or expertise packages.<\/li>\n\n\n\n<li><strong>MCP:<\/strong> connections to external tools and data.<\/li>\n\n\n\n<li><strong>Hooks and commands:<\/strong> automation around defined workflow events.<\/li>\n<\/ul>\n\n\n\n<h2 id=\"pricing\" class=\"wp-block-heading\">Codex vs Claude Code pricing and usage limits<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">The clean consumer comparison begins at $20 per month. Codex access is included through the relevant $20 ChatGPT plan, while Claude Pro is $20 per month in the United States and includes Claude Code. Both offer higher-use consumer tiers at $100 and $200. Similar sticker prices do not mean identical capacity.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Claude Max 5x costs $100 per month and Max 20x costs $200 per month. The \u201c5x\u201d and \u201c20x\u201d names describe usage per session relative to Pro rather than a guaranteed number of coding tasks. Claude and Claude Code share session and weekly limits. Context length, attached files, model choice, and feature use can change how quickly allowance is consumed.<\/p>\n\n\n\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1280\" height=\"800\" src=\"https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/07\/claude-max-plan.webp\" alt=\"Claude Max plan page describing Max 5x and Max 20x usage capacity\" class=\"wp-image-17388\" srcset=\"https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/07\/claude-max-plan.webp 1280w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/07\/claude-max-plan-300x188.webp 300w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/07\/claude-max-plan-1024x640.webp 1024w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/07\/claude-max-plan-768x480.webp 768w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/07\/claude-max-plan-18x12.webp 18w\" sizes=\"(max-width: 1280px) 100vw, 1280px\" \/><figcaption class=\"wp-element-caption\">Claude&#8217;s official Max plan page describes the 5x and 20x usage-capacity options. The monthly prices are stated in the same official article; plan details were checked in July 2026.<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Claude&#8217;s optional usage credits can continue eligible work after included usage when enabled, and those credits are billed separately at standard API rates. API-key authentication is another separate billing route. Calling either route \u201cunlimited\u201d would hide a real marginal cost.<\/p>\n\n\n\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1280\" height=\"800\" src=\"https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/07\/claude-code-plan-limits.webp\" alt=\"Claude Code plan documentation explaining shared session and weekly usage limits\" class=\"wp-image-17389\" srcset=\"https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/07\/claude-code-plan-limits.webp 1280w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/07\/claude-code-plan-limits-300x188.webp 300w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/07\/claude-code-plan-limits-1024x640.webp 1024w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/07\/claude-code-plan-limits-768x480.webp 768w, https:\/\/wp.glbgpt.com\/wp-content\/uploads\/2026\/07\/claude-code-plan-limits-18x12.webp 18w\" sizes=\"(max-width: 1280px) 100vw, 1280px\" \/><figcaption class=\"wp-element-caption\">Claude Code uses the allowance of the connected Pro or Max subscription, including shared session and weekly limits. Exact capacity varies by workload.<\/figcaption><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Codex likewise separates included consumer-plan usage, eligible purchased credits, and API billing. Local messages and cloud chats share a five-hour window, and weekly limits may also apply. A token field in a CLI log cannot be converted directly into a five-hour percentage, subscription credit amount, or API bill.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">For occasional solo coding, the $20 tier on either side is the sensible starting point. Daily repository work with long contexts may justify a higher-use tier, but only after observing real resets and task consumption. Heavy parallel work needs even more care: paying for a larger tier does not remove the need to control context, split tasks, and reserve expensive reasoning for high-judgment work.<\/p>\n\n\n\n<h2 id=\"controlled-test\" class=\"wp-block-heading\">What our controlled Codex vs Claude Code test found<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">We froze four English prompts, a public Python fixture, objective checks, a scoring rubric, and a first-valid-result rule before either agent ran. Tasks covered unfamiliar-repository understanding, a multi-file feature, a DST bug repair, and code review with fixes. Valid results received no human intervention, and weak but valid outputs could not be rerun for a better score.<\/p>\n\n\n\n<style>.comparison-setup-card{--ink:#111318;--body:#414753;--muted:#707784;--line:#e4e7ec;--soft:#f6f7f9;--blue:#315efb;--violet:#7257e8;--green:#087a55;--greenSoft:#eaf8f2;--red:#a23131;--redSoft:#fff0f0;color:var(--ink);font-family:Inter,ui-sans-serif,-apple-system,BlinkMacSystemFont,\"Segoe UI\",sans-serif;line-height:1.5;border:1px solid var(--line);border-radius:8px;background:#fff;overflow:hidden;box-shadow:0 16px 42px rgba(24,32,56,.08)}.comparison-setup-card *{box-sizing:border-box}.comparison-setup-card .s-head{padding:24px 26px 20px;border-bottom:1px solid var(--line)}.comparison-setup-card .s-kicker{margin:0 0 6px;color:var(--muted);font:750 11px\/1.4 ui-monospace,SFMono-Regular,Menlo,monospace;letter-spacing:.08em;text-transform:uppercase}.comparison-setup-card h3{margin:0;font-size:22px;line-height:1.2;letter-spacing:-.025em}.comparison-setup-card .s-sub{margin:7px 0 0;color:var(--body);font-size:13px}.comparison-setup-card .s-body{padding:24px 26px}.comparison-setup-card .s-grid{display:grid;grid-template-columns:repeat(2,1fr);gap:10px}.comparison-setup-card .s-stat{padding:15px;border:1px solid var(--line);border-radius:6px;background:var(--soft)}.comparison-setup-card .s-stat span{display:block;color:var(--muted);font-size:10px;font-weight:800;text-transform:uppercase;letter-spacing:.06em}.comparison-setup-card .s-stat strong{display:block;margin-top:5px;font-size:20px}.comparison-setup-card table{width:100%;border-collapse:collapse;font-size:12px}.comparison-setup-card th,.comparison-setup-card td{padding:11px 12px;border:1px solid var(--line);text-align:left;vertical-align:top}.comparison-setup-card th{background:var(--soft)}.comparison-setup-card .s-note{margin:16px 0 0;padding:12px 14px;border-left:3px solid var(--blue);background:#f3f5ff;color:#38466c;font-size:12px}.comparison-setup-card details{margin-top:12px;border:1px solid var(--line);border-radius:6px}.comparison-setup-card summary{cursor:pointer;padding:12px 14px;font-size:12px;font-weight:800}.comparison-setup-card pre{max-height:560px;overflow:auto;margin:0;padding:14px;white-space:pre-wrap;background:#f5f6f8;color:#222936;font:11px\/1.6 ui-monospace,SFMono-Regular,Menlo,Consolas,monospace}@media(max-width:680px){.comparison-setup-card .s-grid{grid-template-columns:1fr}.comparison-setup-card .s-head,.comparison-setup-card .s-body{padding:20px}}<\/style><aside class=\"comparison-setup-card\"><header class=\"s-head\"><p class=\"s-kicker\">How the test was run<\/p><h3>Same tasks, frozen rules, first valid outputs<\/h3><\/header><div class=\"s-body\"><div class=\"s-grid\"><div class=\"s-stat\"><span>Codex<\/span><strong>gpt-5.6-sol<\/strong><small>Medium T1\u2013T3 \u00b7 High T4<\/small><\/div><div class=\"s-stat\"><span>Claude Code<\/span><strong>claude-opus-5<\/strong><small>Subagent fallback disclosed<\/small><\/div><\/div><table><tbody><tr><th>Fixture<\/th><td>Small public, standard-library-only Python repository<\/td><\/tr><tr><th>Run rule<\/th><td>First valid result preserved; quality reruns prohibited<\/td><\/tr><tr><th>Safety<\/th><td>Fresh copies, offline tasks, no production access, no push or deployment<\/td><\/tr><tr><th>Scoring<\/th><td>Frozen 100-point rubric plus independent public\/hidden checks<\/td><\/tr><\/tbody><\/table><\/div><\/aside>\n\n\n\n<p class=\"wp-block-paragraph\">Codex ran through isolated CLI processes with raw JSONL and CLI-emitted token fields retained. The colleague&#8217;s Claude CLI OAuth session had expired, so Claude Code used a coordinator with fresh subagent contexts. We reconstructed Claude&#8217;s patches in clean audit copies and reran the public and hidden checks. This produced credible practical evidence, but not identical isolation or model-control surfaces.<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table><thead><tr><th>Task<\/th><th>Codex<\/th><th>Claude Code<\/th><th>Bounded interpretation<\/th><\/tr><\/thead><tbody><tr><td>T1: repository understanding<\/td><td>20\/20<\/td><td>20\/20<\/td><td>Tie; compact versus exhaustive<\/td><\/tr><tr><td>T2: multi-file feature<\/td><td>30\/30<\/td><td>30\/30<\/td><td>Functional tie; different validation and test scope<\/td><\/tr><tr><td>T3: DST repair<\/td><td>30\/30<\/td><td>30\/30<\/td><td>Both passed; narrower patch versus broader edge coverage<\/td><\/tr><tr><td>T4: review and repair<\/td><td>18\/20<\/td><td>19\/20<\/td><td>Claude Code showed greater review breadth<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<style>.aggregate-verdict-card{--ink:#111318;--body:#414753;--muted:#707784;--line:#e4e7ec;--soft:#f6f7f9;--blue:#315efb;--violet:#7257e8;--green:#087a55;--greenSoft:#eaf8f2;--red:#a23131;--redSoft:#fff0f0;color:var(--ink);font-family:Inter,ui-sans-serif,-apple-system,BlinkMacSystemFont,\"Segoe UI\",sans-serif;line-height:1.5;border:1px solid var(--line);border-radius:8px;background:#fff;overflow:hidden;box-shadow:0 16px 42px rgba(24,32,56,.08)}.aggregate-verdict-card *{box-sizing:border-box}.aggregate-verdict-card .s-head{padding:24px 26px 20px;border-bottom:1px solid var(--line)}.aggregate-verdict-card .s-kicker{margin:0 0 6px;color:var(--muted);font:750 11px\/1.4 ui-monospace,SFMono-Regular,Menlo,monospace;letter-spacing:.08em;text-transform:uppercase}.aggregate-verdict-card h3{margin:0;font-size:22px;line-height:1.2;letter-spacing:-.025em}.aggregate-verdict-card .s-sub{margin:7px 0 0;color:var(--body);font-size:13px}.aggregate-verdict-card .s-body{padding:24px 26px}.aggregate-verdict-card .s-grid{display:grid;grid-template-columns:repeat(2,1fr);gap:10px}.aggregate-verdict-card .s-stat{padding:15px;border:1px solid var(--line);border-radius:6px;background:var(--soft)}.aggregate-verdict-card .s-stat span{display:block;color:var(--muted);font-size:10px;font-weight:800;text-transform:uppercase;letter-spacing:.06em}.aggregate-verdict-card .s-stat strong{display:block;margin-top:5px;font-size:20px}.aggregate-verdict-card table{width:100%;border-collapse:collapse;font-size:12px}.aggregate-verdict-card th,.aggregate-verdict-card td{padding:11px 12px;border:1px solid var(--line);text-align:left;vertical-align:top}.aggregate-verdict-card th{background:var(--soft)}.aggregate-verdict-card .s-note{margin:16px 0 0;padding:12px 14px;border-left:3px solid var(--blue);background:#f3f5ff;color:#38466c;font-size:12px}.aggregate-verdict-card details{margin-top:12px;border:1px solid var(--line);border-radius:6px}.aggregate-verdict-card summary{cursor:pointer;padding:12px 14px;font-size:12px;font-weight:800}.aggregate-verdict-card pre{max-height:560px;overflow:auto;margin:0;padding:14px;white-space:pre-wrap;background:#f5f6f8;color:#222936;font:11px\/1.6 ui-monospace,SFMono-Regular,Menlo,Consolas,monospace}@media(max-width:680px){.aggregate-verdict-card .s-grid{grid-template-columns:1fr}.aggregate-verdict-card .s-head,.aggregate-verdict-card .s-body{padding:20px}}<\/style><article class=\"aggregate-verdict-card\"><header class=\"s-head\"><p class=\"s-kicker\">Controlled test \u00b7 Frozen rubric<\/p><h3>Codex 98\/100 vs Claude Code 99\/100<\/h3><p class=\"s-sub\">The one-point gap is not a universal winner signal. The useful differences are task-specific.<\/p><\/header><div class=\"s-body\"><div class=\"s-grid\"><div class=\"s-stat\"><span>Codex<\/span><strong>98\/100<\/strong><\/div><div class=\"s-stat\"><span>Claude Code<\/span><strong>99\/100<\/strong><\/div><\/div><table><thead><tr><th>Decision area<\/th><th>Observed result<\/th><th>Practical reading<\/th><\/tr><\/thead><tbody><tr><td>Repository understanding<\/td><td>Tie \u00b7 20\/20 each<\/td><td>Claude was more exhaustive; Codex was more compact.<\/td><\/tr><tr><td>Multi-file feature<\/td><td>Tie \u00b7 30\/30 each<\/td><td>Both passed hidden checks; Claude added broader tests.<\/td><\/tr><tr><td>DST bug repair<\/td><td>Tie \u00b7 30\/30 each<\/td><td>Claude covered more edges; Codex used the narrower patch.<\/td><\/tr><tr><td>Review and repair<\/td><td>Claude 19\/20; Codex 18\/20<\/td><td>Claude found more reproducible defects.<\/td><\/tr><tr><td>Observed speed<\/td><td>Mixed<\/td><td>Claude led T1\/T2; Codex led T3\/T4.<\/td><\/tr><\/tbody><\/table><p class=\"s-note\"><strong>Disclosure:<\/strong> same prompts and fixture, but different isolation paths. Do not generalize this small test to every repository, model tier, or autonomous session.<\/p><\/div><\/article>\n\n\n\n<p class=\"wp-block-paragraph\">The fixture is too small to establish performance on large repositories, every language, every model tier, or long autonomous sessions. The one-point total gap is best read as \u201cboth succeeded, with a review-breadth difference,\u201d not as a universal ranking.<\/p>\n\n\n\n<h2 id=\"user-experience\" class=\"wp-block-heading\">What real users report\u2014and how much to trust it<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Community feedback is valuable when it includes a project, task, model, effort setting, and time horizon. It is weak when a post reports only that one product \u201cfeels smarter\u201d or \u201cis much faster.\u201d Different users compare different repositories, subscription tiers, prompts, extensions, and levels of supervision.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The contextual Reddit example above suggests an interaction trade-off: fast and conversational can demand more attention, while deliberate execution can feel slower but require fewer corrections. Our controlled test only partly overlaps with that account. It found mixed speed and no human intervention in valid runs, so it cannot validate the babysitting claim. This is exactly why community evidence and controlled evidence should be shown together without merging them.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The X harness comment makes a second useful point: a coding result belongs to the whole system, not just the model label. Neither individual post should be treated as survey data. Use them to identify questions for your own trial\u2014intervention count, diff size, test quality, and recovery from a wrong assumption.<\/p>\n\n\n\n<h2 id=\"globalgpt\" class=\"wp-block-heading\">Extending Codex and Claude Code with GlobalGPT<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">GlobalGPT is a <a href=\"https:\/\/www.glbgpt.com\/hub\/best-all-in-one-ai-tools-and-platforms-11-options-compared\/\">multi-model, multimodal workspace<\/a>, not a replacement for native repository agents. Its practical role is to extend an existing host: use another available agent model for a second opinion, plan review, documentation pass, or specialized output while Codex or Claude Code remains responsible for repository access, shell commands, tests, and approvals.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">GlobalGPT publishes first-party CLI guides for <a href=\"https:\/\/www.glbgpt.com\/hub\/how-to-use-globalgpt-cli-in-codex\/\">Codex<\/a> and <a href=\"https:\/\/www.glbgpt.com\/hub\/how-to-use-globalgpt-cli-in-claude-code\/\">Claude Code<\/a>, as well as Cursor. In our safe T5 integration test, the CLI completed the same frozen task in both compared environments. In Codex, we also verified a read-only MCP model-list call and installed, loaded, and used the GlobalGPT Skill.<\/p>\n\n\n\n<style>.globalgpt-integration-card{--ink:#111318;--body:#414753;--muted:#707784;--line:#e4e7ec;--soft:#f6f7f9;--blue:#315efb;--violet:#7257e8;--green:#087a55;--greenSoft:#eaf8f2;--red:#a23131;--redSoft:#fff0f0;color:var(--ink);font-family:Inter,ui-sans-serif,-apple-system,BlinkMacSystemFont,\"Segoe UI\",sans-serif;line-height:1.5;border:1px solid var(--line);border-radius:8px;background:#fff;overflow:hidden;box-shadow:0 16px 42px rgba(24,32,56,.08)}.globalgpt-integration-card *{box-sizing:border-box}.globalgpt-integration-card .s-head{padding:24px 26px 20px;border-bottom:1px solid var(--line)}.globalgpt-integration-card .s-kicker{margin:0 0 6px;color:var(--muted);font:750 11px\/1.4 ui-monospace,SFMono-Regular,Menlo,monospace;letter-spacing:.08em;text-transform:uppercase}.globalgpt-integration-card h3{margin:0;font-size:22px;line-height:1.2;letter-spacing:-.025em}.globalgpt-integration-card .s-sub{margin:7px 0 0;color:var(--body);font-size:13px}.globalgpt-integration-card .s-body{padding:24px 26px}.globalgpt-integration-card .s-grid{display:grid;grid-template-columns:repeat(2,1fr);gap:10px}.globalgpt-integration-card .s-stat{padding:15px;border:1px solid var(--line);border-radius:6px;background:var(--soft)}.globalgpt-integration-card .s-stat span{display:block;color:var(--muted);font-size:10px;font-weight:800;text-transform:uppercase;letter-spacing:.06em}.globalgpt-integration-card .s-stat strong{display:block;margin-top:5px;font-size:20px}.globalgpt-integration-card table{width:100%;border-collapse:collapse;font-size:12px}.globalgpt-integration-card th,.globalgpt-integration-card td{padding:11px 12px;border:1px solid var(--line);text-align:left;vertical-align:top}.globalgpt-integration-card th{background:var(--soft)}.globalgpt-integration-card .s-note{margin:16px 0 0;padding:12px 14px;border-left:3px solid var(--blue);background:#f3f5ff;color:#38466c;font-size:12px}.globalgpt-integration-card details{margin-top:12px;border:1px solid var(--line);border-radius:6px}.globalgpt-integration-card summary{cursor:pointer;padding:12px 14px;font-size:12px;font-weight:800}.globalgpt-integration-card pre{max-height:560px;overflow:auto;margin:0;padding:14px;white-space:pre-wrap;background:#f5f6f8;color:#222936;font:11px\/1.6 ui-monospace,SFMono-Regular,Menlo,Consolas,monospace}@media(max-width:680px){.globalgpt-integration-card .s-grid{grid-template-columns:1fr}.globalgpt-integration-card .s-head,.globalgpt-integration-card .s-body{padding:20px}}<\/style><article class=\"globalgpt-integration-card\"><header class=\"s-head\"><p class=\"s-kicker\">GlobalGPT integration \u00b7 Excluded from coding score<\/p><h3>CLI verified in both workflows; MCP and Skill verified in Codex<\/h3><p class=\"s-sub\">This measures integration paths, not whether GlobalGPT replaces either native coding agent.<\/p><\/header><div class=\"s-body\"><table><thead><tr><th>Path<\/th><th>Codex environment<\/th><th>Claude colleague run<\/th><\/tr><\/thead><tbody><tr><td>GlobalGPT CLI<\/td><td>Verified with gpt-5.6-sol and gpt-5.6-luna<\/td><td>Verified with the same frozen tasks<\/td><\/tr><tr><td>MCP<\/td><td>Interactive read-only model list verified<\/td><td>Not connected during the run<\/td><\/tr><tr><td>Skill<\/td><td>Installed, loaded, and exercised<\/td><td>Installed but not used for frozen calls<\/td><\/tr><tr><td>Unattended MCP<\/td><td>Approval limitation observed<\/td><td>Not tested<\/td><\/tr><\/tbody><\/table><details><summary>View full integration evidence<\/summary><pre># GlobalGPT Integration Test Summary\n\n## Scope\n\nT5 measures GlobalGPT integration and is excluded from the native Codex coding score. Yukie was not used.\n\n## Model lock and readiness\n\n- Model A: `gpt-5.6-sol`\n- Model B: `gpt-5.6-luna`\n- Both models appeared in the live model list.\n- Both lightweight readiness probes returned valid parseable text.\n- Models were locked before formal T5A and T5B output calls.\n\n## Formal CLI results\n\n| Task | Model | Status | Duration | Prompt tokens | Completion tokens | Required headings |\n|---|---|---|---:|---:|---:|---:|\n| T5A | gpt-5.6-sol | Valid | 27 s | 2,116 | 1,207 | 6\/6 |\n| T5B | gpt-5.6-luna | Valid | 14 s | 2,116 | 1,441 | 6\/6 |\n\nBoth formal calls used the same frozen English product-planning task through `glbgpt exec`, returned English-only visible output, and exposed no credentials. No quality rerun occurred.\n\n## MCP result\n\n- `codex mcp list` showed `globalgpt` enabled.\n- A non-interactive Codex session discovered and attempted `mcp__globalgpt__glbgpt_list_models`, but the call was cancelled because unattended MCP approval was unavailable. No safety setting was relaxed.\n- The active interactive Codex session successfully called the same read-only GlobalGPT MCP tool and received nine chat models.\n- Result: interactive MCP access verified; unattended non-interactive MCP execution retains an approval limitation.\n\n## Skill result\n\n- The GlobalGPT and GlobalGPT Coding skills are installed in the Codex skill directory through links to the bundled CLI skill package.\n- The GlobalGPT skill was loaded and followed for configured-session checking, live model selection, readiness, credit-safe options, and failure handling.\n- Result: installed and exercised in the active Codex workflow.\n\n## Bounded conclusion\n\nThis run verifies working GlobalGPT CLI calls, an interactive GlobalGPT MCP read-only call from Codex, and an installed\/used GlobalGPT Skill path. It does not prove that every model, media tool, host, account, or unattended MCP configuration works. It does not test Yukie or support a claim that GlobalGPT replaces Codex or Claude Code.\n<\/pre><\/details><p class=\"s-note\">Verified scope: specific CLI calls, one interactive MCP read, and one installed\/used Skill path. This does not prove every model, media tool, host, or unattended configuration.<\/p><\/div><\/article>\n\n\n\n<p class=\"wp-block-paragraph\">The boundary matters. The colleague&#8217;s Claude run verified the CLI path, not Claude-side MCP. The Skill was installed there but was not used for the frozen calls. An unattended Codex MCP attempt still needed manual approval. Model availability and credits also depend on the current GlobalGPT plan, so developers should check the live list rather than assume every model is included.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">GlobalGPT also covers image, video, audio, and guided Agent workflows. Yukie&#8217;s Slides, Document, and Image entries are a different category: they offer explicit templates and step-by-step browser guidance for non-code tasks. That can be easier than opening a coding CLI to make a presentation or structured document, but it does not replace repository reading, test execution, or code review.<\/p>\n\n\n\n<style>.yukie-workflow-summary-card{--ink:#111318;--body:#414753;--muted:#707784;--line:#e4e7ec;--soft:#f6f7f9;--blue:#315efb;--violet:#7257e8;--green:#087a55;--greenSoft:#eaf8f2;--red:#a23131;--redSoft:#fff0f0;color:var(--ink);font-family:Inter,ui-sans-serif,-apple-system,BlinkMacSystemFont,\"Segoe UI\",sans-serif;line-height:1.5;border:1px solid var(--line);border-radius:8px;background:#fff;overflow:hidden;box-shadow:0 16px 42px rgba(24,32,56,.08)}.yukie-workflow-summary-card *{box-sizing:border-box}.yukie-workflow-summary-card .s-head{padding:24px 26px 20px;border-bottom:1px solid var(--line)}.yukie-workflow-summary-card .s-kicker{margin:0 0 6px;color:var(--muted);font:750 11px\/1.4 ui-monospace,SFMono-Regular,Menlo,monospace;letter-spacing:.08em;text-transform:uppercase}.yukie-workflow-summary-card h3{margin:0;font-size:22px;line-height:1.2;letter-spacing:-.025em}.yukie-workflow-summary-card .s-sub{margin:7px 0 0;color:var(--body);font-size:13px}.yukie-workflow-summary-card .s-body{padding:24px 26px}.yukie-workflow-summary-card .s-grid{display:grid;grid-template-columns:repeat(2,1fr);gap:10px}.yukie-workflow-summary-card .s-stat{padding:15px;border:1px solid var(--line);border-radius:6px;background:var(--soft)}.yukie-workflow-summary-card .s-stat span{display:block;color:var(--muted);font-size:10px;font-weight:800;text-transform:uppercase;letter-spacing:.06em}.yukie-workflow-summary-card .s-stat strong{display:block;margin-top:5px;font-size:20px}.yukie-workflow-summary-card table{width:100%;border-collapse:collapse;font-size:12px}.yukie-workflow-summary-card th,.yukie-workflow-summary-card td{padding:11px 12px;border:1px solid var(--line);text-align:left;vertical-align:top}.yukie-workflow-summary-card th{background:var(--soft)}.yukie-workflow-summary-card .s-note{margin:16px 0 0;padding:12px 14px;border-left:3px solid var(--blue);background:#f3f5ff;color:#38466c;font-size:12px}.yukie-workflow-summary-card details{margin-top:12px;border:1px solid var(--line);border-radius:6px}.yukie-workflow-summary-card summary{cursor:pointer;padding:12px 14px;font-size:12px;font-weight:800}.yukie-workflow-summary-card pre{max-height:560px;overflow:auto;margin:0;padding:14px;white-space:pre-wrap;background:#f5f6f8;color:#222936;font:11px\/1.6 ui-monospace,SFMono-Regular,Menlo,Consolas,monospace}@media(max-width:680px){.yukie-workflow-summary-card .s-grid{grid-template-columns:1fr}.yukie-workflow-summary-card .s-head,.yukie-workflow-summary-card .s-body{padding:20px}}<\/style><article class=\"yukie-workflow-summary-card\"><header class=\"s-head\"><p class=\"s-kicker\">Different workflow category<\/p><h3>Yukie guided browser-agent tests<\/h3><p class=\"s-sub\">Useful for task-specific web workflows; not tested as a repository coding-agent replacement.<\/p><\/header><div class=\"s-body\"><table><thead><tr><th>Workflow<\/th><th>Criteria completed<\/th><th>First-valid result<\/th><\/tr><\/thead><tbody><tr><td>Slides<\/td><td>5\/6<\/td><td>Five-slide deck generated; speaker notes could not be verified.<\/td><\/tr><tr><td>Document<\/td><td>6\/6<\/td><td>Structured research brief with export and version history.<\/td><\/tr><tr><td>Image + caption<\/td><td>5\/6<\/td><td>Strong square visual; required caption was missing.<\/td><\/tr><\/tbody><\/table><p class=\"s-note\">Safe conclusion: clear entry points and guided multimodal workflows. Unsupported conclusions: unlimited usage, exact price, or full replacement of Codex\/Claude Code.<\/p><\/div><\/article>\n\n\n\n<h2 id=\"which-one\" class=\"wp-block-heading\">Which coding agent should an independent developer choose?<\/h2>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>Choose Codex<\/strong> when compact bounded execution, its combination of available surfaces, or the run evidence available in your setup fits how you work.<\/li>\n\n\n\n<li><strong>Choose Claude Code<\/strong> when a conversational terminal loop, the customization features you have verified, or broader review behavior like that observed in our fixture matters more.<\/li>\n\n\n\n<li><strong>Use either for an unfamiliar repository<\/strong> only after a read-only architecture and risk pass. Both tools handled that task well.<\/li>\n\n\n\n<li><strong>Use both selectively<\/strong> when a second-agent review is worth the extra subscription and coordination cost. Ask the second tool to challenge a diff or plan rather than blindly repeat the task.<\/li>\n\n\n\n<li><strong>Add GlobalGPT<\/strong> when multi-model routing or multimodal work is the missing layer. Keep native repository operations in the coding host.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">A seven-day trial is more informative than a leaderboard. Pick one real but non-production feature, bug, and review. Freeze the prompt and success command. Record the first valid diff, tests passed, elapsed time, human interventions, unexpected scope, and how the work affected your available usage. The better product is the one whose mistakes you can detect and whose workflow you can sustain.<\/p>\n\n\n\n<h2 id=\"faq\" class=\"wp-block-heading\">Codex vs Claude Code FAQ<\/h2>\n\n\n\n<h3 class=\"wp-block-heading\">Is Codex better than Claude Code?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Not universally. Both completed the four tasks in our fixture. Claude Code led in review breadth, while Codex was more compact in parts of implementation and bug repair.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Which is better for large repositories?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Our small fixture cannot answer that. Evaluate both on a read-only architecture task from your own repository before allowing changes.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Which coding agent requires less supervision?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">It depends on task clarity, model settings, repository instructions, and the desired interaction style. One Reddit user reported different supervision patterns, but that individual experience is not a product-wide rule.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Which is cheaper?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">The main consumer tiers align at $20, $100, and $200 per month, but included capacity and limit mechanics differ. Optional credits and API billing are separate.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Do Claude and Claude Code share usage limits?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Yes. Claude and Claude Code use shared session and weekly limits on connected Pro or Max plans.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Do Codex local and cloud tasks share limits?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Local messages and cloud chats share a five-hour window, and weekly limits may also apply.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Can GlobalGPT replace Codex or Claude Code?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">No. It can extend either workflow with additional models, CLI, MCP, Skills, and multimodal tools, but repository access and coding-agent behavior remain in the host.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Can I use GlobalGPT from both products?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">GlobalGPT provides official CLI tutorials for both. We verified CLI use in both environments and MCP plus Skill use in Codex, with an approval limitation for unattended MCP.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><em>Pricing, plan rules, integrations, and source evidence were checked in July 2026. Product details can change.<\/em><\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><a href=\"https:\/\/www.glbgpt.com\/home?inviter=hub_content_home&amp;login=1\">Open GlobalGPT<\/a> if you want to add a multi-model and multimodal layer to your existing coding workflow.<\/p>\n\n\n<\/article>","protected":false},"excerpt":{"rendered":"<p>Codex vs Claude Code is less about finding one universal winner and more about choosing the working style you can trust every day. Both products can inspect a repository, edit multiple files, run commands, execute tests, and explain a patch. The meaningful differences appear in how you delegate work, how often you steer the agent, [&hellip;]<\/p>\n","protected":false},"author":7,"featured_media":17361,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"_seopress_robots_primary_cat":"","_seopress_titles_title":" Codex vs Claude Code: Features, Pricing & Hands-On Tests","_seopress_titles_desc":"A neutral Codex vs Claude Code comparison covering features, pricing, limits, real repository tests, user experience, and GlobalGPT integrations.","_seopress_robots_index":"","footnotes":""},"categories":[7],"tags":[],"class_list":["post-17342","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai-chat"],"_links":{"self":[{"href":"https:\/\/wp.glbgpt.com\/wp-json\/wp\/v2\/posts\/17342","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/wp.glbgpt.com\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/wp.glbgpt.com\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/wp.glbgpt.com\/wp-json\/wp\/v2\/users\/7"}],"replies":[{"embeddable":true,"href":"https:\/\/wp.glbgpt.com\/wp-json\/wp\/v2\/comments?post=17342"}],"version-history":[{"count":2,"href":"https:\/\/wp.glbgpt.com\/wp-json\/wp\/v2\/posts\/17342\/revisions"}],"predecessor-version":[{"id":17393,"href":"https:\/\/wp.glbgpt.com\/wp-json\/wp\/v2\/posts\/17342\/revisions\/17393"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/wp.glbgpt.com\/wp-json\/wp\/v2\/media\/17361"}],"wp:attachment":[{"href":"https:\/\/wp.glbgpt.com\/wp-json\/wp\/v2\/media?parent=17342"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/wp.glbgpt.com\/wp-json\/wp\/v2\/categories?post=17342"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/wp.glbgpt.com\/wp-json\/wp\/v2\/tags?post=17342"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}