| Name | Modified | Size | Downloads / Week |
|---|---|---|---|
| Parent folder | |||
| README.md | 2026-09-11 | 11.9 kB | |
| v0.9.8 source code.tar.gz | 2026-09-11 | 864.5 kB | |
| v0.9.8 source code.zip | 2026-09-11 | 1.2 MB | |
| Totals: 3 Items | 2.0 MB | 0 | |
13 commits since v0.9.7, closing the 14 issues of the v0.9.8 milestone.
v0.9.7 made the block layer able to report a failed transfer. This release is what happened next: a way to cause that failure on demand, and then the error paths it reached. Five of the fourteen issues are code that had only ever been verified by inspection, because the condition it handles could not be produced — and four of those five turned out to be wrong.
The second half of the release is the format string: a kernel printf that executed user input, that could not print a 64-bit value, and that nothing checked.
Failures on demand
/proc/faultinj arms a failure for the next block transfer (#338):
echo "read 1" > /proc/faultinj # fail the next sector read
echo "read 3 sector 40 skip 1" > /proc/faultinj # fail reads of sector 40, after letting one through
echo "off" > /proc/faultinj # disarm
cat /proc/faultinj # what is armed, what has fired, what was read last
It sits inside ata_device_read_sector_pio before the transfer, which is where every real error path in that function leaves off, so the code above sees exactly what a failing device shows it. Compiled in only under -DENABLE_ATA_FAULT_INJECTION=ON; the default build keeps nothing — size on the object reports 0 text, 0 data, 0 bss. The entry is mode 0600, because leaving it armed is a way to lose data.
Two things it needed before it could reach anything interesting, and both are the point: a sector target, because a bare count of failures is spent by path resolution long before the read under test, and a skip count, because every path that writes a metadata block reads something from the same block first through a function that already checks its read.
What it found
A failed metadata read wrote a zeroed block back (#344, [#342], [#343]). Several ext2 functions read a metadata block into a cache, patch one field and write the whole block back, and none checked the read. ext2_alloc_cache zeroes what it returns and ata_read fills only up to the failing sector, so a failed read leaves the block correct up to that point and zeroes after it. Writing that back destroys the rest.
With the inode-table read check reverted and one sector made to fail:
not ok 71 - t_meta_rmw: Exit: 1
[t_meta_rmw] /home/user/t_rmw_1.bin is 0 bytes, expected 7: its inode was overwritten
[t_meta_rmw] /home/user/t_rmw_1.bin has mode 0: its inode was zeroed
... seven neighbours, all zeroed
Twenty-two call sites discarded a read or write result; all of them check now. The bitmap case was the worst: an all-zero bitmap written back marks every block in the group free, and the allocator then hands out storage that is in use. The governing rule is narrow — never write a metadata block you could not read. A block that cannot be freed is a leak; a group wrongly marked free is corruption.
A block whose mapping could not be read was reported as a hole (#356). ext2_get_real_block_index returned the block on the device as its value, and 0 meant two opposite things: the file has a hole there and must read as zeros, or the index block holding the pointer could not be read. read() took the first reading, so a file with an unreadable index block came back as a sparse file — a buffer of zeros and success.
The mapping returns 0/-errno now, with the index leaving through a parameter. Seventeen callers of ext2_read_block and friends tested for == -1 and so swallowed the -EIO the block layer has returned since [#291]; without fixing those the new error could not reach userspace, which is what the first run of the test caught.
An I/O error during path resolution was read as "not a symbolic link" (#353). __is_a_link answered 0 whenever vfs_stat failed, and vfs_stat fails both for a component that is not there and for one that could not be read. A link that could not be read was resolved as an ordinary name, and the walk carried on along the path the link would have redirected away from. Only -ENOENT is an answer now. The reason had to be made available first: ext2_find_direntry reported -1 for a NULL argument, a parent that is not a directory, an unreadable inode and an absent entry alike, and ext2_stat turned all four into -ENOENT.
Memory safety in ext2 directories
A directory entry was written past the end of its block (#360). ext2_append_new_direntry shrinks the last entry of a block to its real length and puts the new one after it, but measured the free space from the start of that last entry instead of its end — so the check passed with real_rec_len bytes less room than it thought. The name went over whatever the slab allocator kept next to the block buffer, and never reached the disk, so the entry described a name that was not there.
It was silent: three events per suite run on develop, 7 and 255 bytes past the end. It became a 100%-deterministic kernel panic the moment an unrelated change to the allocation pattern left the neighbouring buffer on a free list, which is how it was found — the panic pointed at the slab allocator, and the slab allocator was the victim. ext2_initialize_direntry now refuses a record the name does not fit in, and all four call sites check it.
creat() on a directory returned a writable descriptor for it (#346). It now fails with EISDIR, and creat() on an existing regular file truncates it, which it never did.
The format string
sys_syslog executed the user's message as a kernel format string (#351). The message went to dbg_printf as its format argument with no variadic arguments after it, and the syscall is reachable by any unprivileged program. Run deliberately, from an ordinary test program: %s printed kernel code and stack bytes to the console, %p and %x printed kernel addresses, and a wide field width ended the run at the emulator timeout with 35 of 69 tests completed. The fourth was %n, which the kernel's printf implemented as a write through a pointer taken from the same nonexistent argument list.
The message is now an argument with "%s" as the format, and %n is gone from both printf implementations — dbg_printf calls the klib one, so patching only the libc copy would have left the primitive intact.
%lld printed literally and misaligned everything after it (#332). The length modifiers were parsed but every argument was read through a four-byte va_arg. The number paths carry unsigned long long end to end now, and because a -nostdlib libc cannot call the compiler's 64-bit division helpers, the digit conversion divides with shifts and subtracts, with a native fast path for values that fit in 32 bits.
Nothing checked a format against its arguments (#216). dbg_printf, __syslog and the whole printf family now carry __attribute__((format(printf, ...))), so a mismatch is a compile-time error under the existing -Werror. It immediately found four live bugs — five vfs_* diagnostics logging %s with no argument at all, printing stack contents as the path — and about 120 cases of type drift.
...but only in the calls a build actually compiles (#365). pr_debug and friends compile to nothing below each file's __DEBUG_LEVEL__, which is LOGLEVEL_NOTICE almost everywhere, so most log calls in the tree were never checked by [#216] or by anything else. Building with every level forced on found 85 further mismatches, hidden below the threshold — including a pr_debug with a missing argument in both copies of multiboot.c, one with an argument too many in ext2_debug.c, and a strsignal() result passed straight to %s in sys_kill where it can be NULL. A FORCE_DEBUG_LOGLEVEL build option and a CI job now keep it that way; the job runs in 25 seconds.
size_t was not the compiler's size_t (#363). lib/inc/stddef.h typedef'd it to unsigned long, which is the same width as the compiler's own unsigned int on i686 but not the same type, and -Wformat compares types. Every %zu in the tree was therefore a warning, and the tree had been paying for it one hand-written %lu at a time. Now:
:::c
typedef __SIZE_TYPE__ size_t;
typedef __PTRDIFF_TYPE__ ssize_t;
%zu and %zd are correct everywhere, on both the Linux and the Homebrew i686 toolchains, and 51 sites that had been written around the problem were put back.
Verification
75 tests, 0 failures — two consecutive Debug runs, one Release run, and one Debug run with -DENABLE_ATA_FAULT_INJECTION=ON, all on the tagged tree. Release was checked separately because -Wformat-overflow only runs at -O2, and it is one of the checks the new CI job depends on.
Seven tests were added: t_faultinj, t_meta_rmw, t_syslog_format, t_printf_length, t_dir_block_boundary, t_indirect_map, t_namei_io. The suite went from 68 to 75.
Every fix carries a negative control: the fix reverted, the specific test shown failing with its exact output, then restored. Two of them read:
[t_indirect_map] the read of an unreadable index block returned 16 bytes, all zeros: the file was reported as sparse
[t_namei_io] opening /home/user/t_namei_io.txt succeeded while the root directory could not be read: the error was read as `not a symbolic link` and the resolution carried on
Known issues
Not regressions. Each has a reproduction in its issue, and the four filed during this cycle are the residue of reading the same code closely.
- #372 a directory block that cannot be read ends the iterator walk, and every caller reads that as the end of the directory.
ext2_directory_is_emptytherefore answers "empty", and its only caller isrmdir— so a directory whose first block cannot be read is removed and its contents orphaned. Same outcome as [#341], different cause. The most serious thing known to be open. - #373 twelve callers of
ext2_alloc_cachestill do not check for NULL; five are plain null dereferences, andext2_getdentsreturns an empty listing and success. - #371
ext2_readlinktakes the target length fromstrlenover a field ext2 does not terminate, and has no branch for a target held in a data block. - #288 the symlink substitution in
__resolve_pathwrites without a terminator and without a bound. Filed long ago as unreachable; the image does in fact contain three symlinks, so the branch runs on every boot that touches them — it survives on all three targets being longer than the name they replace, not on a check. - #191 / #259 syscalls do not validate user pointers. A deliberate state of a teaching kernel, but not a property to assume away.
- #333 the ATA 48-bit offset is truncated by every caller.
- #336 intermittent single-test failures that have not reproduced since [#291].
Corrections made during this cycle
- The regression test for [#360] does not fail on the unfixed tree, and its header says so. The out-of-bounds write landed in a neighbouring buffer that kept the bytes, so every lookup still found the name; the deterministic pre-fix failure was the panic, under an allocation pattern the test cannot arrange. The detector is the bound check itself, not the test.
- The first fix for [#356] was incomplete in a way its own test caught: the new
-EIOwas dropped by an== -1comparison two frames up and the zeroed buffer was copied out anyway. Seventeen such comparisons are part of the fix. - I filed the
__resolve_pathsymlink defects as a new issue before finding [#288], which had them both, three years of context, and a reachability analysis that is now out of date. The duplicate is closed and the correction is on [#288].