dar-support Mailing List for DAR - Disk ARchive
For full, incremental, compressed and encrypted backups or archives
Brought to you by:
edrusb
You can subscribe to this list here.
| 2003 |
Jan
|
Feb
|
Mar
|
Apr
|
May
|
Jun
|
Jul
|
Aug
|
Sep
|
Oct
(1) |
Nov
|
Dec
|
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 2004 |
Jan
|
Feb
(21) |
Mar
(37) |
Apr
(8) |
May
(23) |
Jun
(13) |
Jul
(41) |
Aug
(12) |
Sep
(58) |
Oct
(13) |
Nov
(34) |
Dec
(17) |
| 2005 |
Jan
(49) |
Feb
(98) |
Mar
(33) |
Apr
(41) |
May
(48) |
Jun
(24) |
Jul
(45) |
Aug
(25) |
Sep
(22) |
Oct
(26) |
Nov
(60) |
Dec
(28) |
| 2006 |
Jan
(63) |
Feb
(45) |
Mar
(29) |
Apr
(44) |
May
(19) |
Jun
(8) |
Jul
(32) |
Aug
(36) |
Sep
(24) |
Oct
(61) |
Nov
(84) |
Dec
(93) |
| 2007 |
Jan
(77) |
Feb
(41) |
Mar
(24) |
Apr
(32) |
May
(25) |
Jun
(36) |
Jul
(70) |
Aug
(21) |
Sep
(37) |
Oct
(18) |
Nov
(23) |
Dec
(6) |
| 2008 |
Jan
(9) |
Feb
(13) |
Mar
(8) |
Apr
(4) |
May
|
Jun
(4) |
Jul
(21) |
Aug
(4) |
Sep
(8) |
Oct
(29) |
Nov
(24) |
Dec
(16) |
| 2009 |
Jan
(13) |
Feb
(33) |
Mar
(20) |
Apr
(21) |
May
(22) |
Jun
(5) |
Jul
(40) |
Aug
(2) |
Sep
(2) |
Oct
(10) |
Nov
(22) |
Dec
(13) |
| 2010 |
Jan
(2) |
Feb
(9) |
Mar
(13) |
Apr
(15) |
May
(26) |
Jun
(3) |
Jul
(10) |
Aug
(7) |
Sep
(5) |
Oct
(21) |
Nov
(4) |
Dec
(17) |
| 2011 |
Jan
(22) |
Feb
(23) |
Mar
(22) |
Apr
(12) |
May
|
Jun
(39) |
Jul
(16) |
Aug
(7) |
Sep
(4) |
Oct
|
Nov
(19) |
Dec
(11) |
| 2012 |
Jan
(101) |
Feb
(5) |
Mar
(18) |
Apr
(9) |
May
(3) |
Jun
(27) |
Jul
(17) |
Aug
(19) |
Sep
(4) |
Oct
(30) |
Nov
(12) |
Dec
(23) |
| 2013 |
Jan
(14) |
Feb
(5) |
Mar
(26) |
Apr
(17) |
May
(18) |
Jun
(28) |
Jul
(12) |
Aug
(11) |
Sep
(5) |
Oct
(24) |
Nov
(9) |
Dec
(1) |
| 2014 |
Jan
(29) |
Feb
(19) |
Mar
(4) |
Apr
(9) |
May
(2) |
Jun
(1) |
Jul
|
Aug
(11) |
Sep
(10) |
Oct
(3) |
Nov
(25) |
Dec
(6) |
| 2015 |
Jan
(5) |
Feb
(15) |
Mar
(5) |
Apr
(15) |
May
(9) |
Jun
(15) |
Jul
(13) |
Aug
(3) |
Sep
(33) |
Oct
(32) |
Nov
(10) |
Dec
|
| 2016 |
Jan
(11) |
Feb
(24) |
Mar
(4) |
Apr
(41) |
May
(7) |
Jun
(28) |
Jul
(17) |
Aug
(4) |
Sep
(4) |
Oct
(3) |
Nov
(9) |
Dec
(24) |
| 2017 |
Jan
(27) |
Feb
(20) |
Mar
(19) |
Apr
(24) |
May
(9) |
Jun
(5) |
Jul
(16) |
Aug
(5) |
Sep
(28) |
Oct
|
Nov
(7) |
Dec
(12) |
| 2018 |
Jan
(4) |
Feb
(10) |
Mar
(11) |
Apr
(2) |
May
|
Jun
|
Jul
(25) |
Aug
(5) |
Sep
(29) |
Oct
(11) |
Nov
(6) |
Dec
(16) |
| 2019 |
Jan
(12) |
Feb
(35) |
Mar
(1) |
Apr
(2) |
May
(31) |
Jun
(12) |
Jul
(14) |
Aug
(40) |
Sep
|
Oct
(20) |
Nov
|
Dec
(8) |
| 2020 |
Jan
(37) |
Feb
(34) |
Mar
|
Apr
(6) |
May
(24) |
Jun
(7) |
Jul
(13) |
Aug
|
Sep
|
Oct
|
Nov
(6) |
Dec
|
| 2021 |
Jan
(17) |
Feb
(22) |
Mar
(10) |
Apr
(54) |
May
(40) |
Jun
|
Jul
(20) |
Aug
(10) |
Sep
(7) |
Oct
(10) |
Nov
(11) |
Dec
(30) |
| 2022 |
Jan
(11) |
Feb
(9) |
Mar
|
Apr
(7) |
May
(22) |
Jun
(19) |
Jul
(8) |
Aug
(6) |
Sep
(7) |
Oct
(5) |
Nov
(11) |
Dec
|
| 2023 |
Jan
(1) |
Feb
(2) |
Mar
(13) |
Apr
|
May
(3) |
Jun
(42) |
Jul
(19) |
Aug
(15) |
Sep
(21) |
Oct
|
Nov
(12) |
Dec
(33) |
| 2024 |
Jan
(4) |
Feb
(4) |
Mar
|
Apr
(20) |
May
(4) |
Jun
(2) |
Jul
|
Aug
|
Sep
(3) |
Oct
|
Nov
|
Dec
(3) |
| 2025 |
Jan
(3) |
Feb
(8) |
Mar
(26) |
Apr
(10) |
May
(5) |
Jun
(7) |
Jul
|
Aug
(2) |
Sep
|
Oct
(2) |
Nov
(3) |
Dec
(21) |
| 2026 |
Jan
(8) |
Feb
(1) |
Mar
(3) |
Apr
(7) |
May
(2) |
Jun
(12) |
Jul
|
Aug
|
Sep
|
Oct
|
Nov
|
Dec
|
|
From: mannino <ma...@of...> - 2026-06-20 14:04:55
|
>> Just before the shutdown, one of those systems has created a full dar >> archive with the same kind of corruption (direct mode doesn't work, -t >> -0 reports 5 CRC errors). After powering it on and manually running >> full dar backup, a valid archive was produced. > > this is quite weird... we get close to the invocation of the cosmic > particle that have changed something somewhere in some memory resident > code... (I'm almost kidding!) I actually thought about that but for that to happen at least 3 times in 2 different systems/locations and time... >> 2. Byte sizes of non/last slices would allow basic integrity check >> but >> (1) `file` doesn't show them; (2) dar doesn't show them in -l -0 mode >> (so unable to extract from invalid file); (3) dar -l -q works on >> isolated catalogue but only reports non-last slice size - given our >> problem here is likely last slice being truncated, it doesn't really >> help. > because this information is not stored inside the archive data, but > found in the metadata (here the filenames) with in addition the side > effect of code modularity: Thanks for the explanation. You have previously said that slice-layout holds slice sizes and apparently it's considered part of "archive data"... > The sar layer (which manages slices) > processes a stream of bytes from upper layers (ciphering, compression, > filtering, and so forth) which is generated on the fly from the > filesystem to backup, the sar takes this flow and creates slices from it > up to the time the flow dries up. > > What is stored at the beginning of *each slice* is the so called > slice-layout: > - first slice max size, > - other slices max size, > - slice header size > - and slice format (2 formats so far). > > Only this could be shown in sequential-read mode, not what is displayed > today in direct mode which adds to this the overall size of the archive > and the size of the last slice (easy to get because in that mode this is > the last slice that get read first). Do I understand it correctly that the abstraction doesn't allow reading/provide access to slice-layout data, making it impossible to view slice header in -l -q -0 (or any other) mode? Isn't at least the first slice's header exposed? Or the main concern here is output verbosity (see next)? >>>> Okay, the ability to view the details of slice-layout in -0 mode would >>>> be welcome. I'd say any header info available to dar should be made >>>> visible on equal terms. >>> >>> I guess when you drive a car, you don't seen all the internal sensor >>> values (there is a lot of them in today's car), but over the classical >>> few indicators (oil temperature, speed, motor RPM,...) this is only when >>> something wrong happens that these internal sensors trigger an alarm on >>> your car dashboard in human understandable form, no? >> >> This is why I'm always adding "in -0 mode". This mode is exactly for >> "when something wrong happens", right? > > No. The sequential-read mode is mainly to cope with devices that do not > provide direct access mode to a given offset of a file. [...] > Today it is discouraged to remove tape marks while this would speed up > the backup and reading processes, because tape marks (with the metadata > that follows) can also be used to rebuild the internal catalogue of a > truncated archive (-y option), it provides saved files metadata > redundancy within a given archive. Okay, -0 is "mainly" for unseekable devices but you state that tape marks are recommended even in normal backups for fault tolerance/redundancy meaning that -0's secondary purpose is still "when something wrong happens". I think I've also seen in the docs that for recovery, you're expected to try -0 first, then -al, then -al -0 as last resort. In this case, provided that the abstraction permits, why not make extra info (slice sizes here) available to the user in -0 mode? (It's useless for -al since user has to supply headers' data himself.) dar already provides a dozen of -v(erbose) options, if output is a concern then it may be brought under -vm or something. > I've added the following in the TODO list (at sourceforge) : expose the > internal_name and data_name a given archive in that -l -q mode (or -l > -v). This will be a per archive information and not per slice, as dar > does not handle slices individually. Thanks, I'm sure this will be really handy sometimes. > I think your homebrew system is missing metadata support. Without it, > you are stuck to the them same approach as tapes and all the > restrictions it has. In fact, your homebrew system is probably even more > restrictive than a tape system, as users usually add a sticker on their > tapes when manually handled or the robot is hopefully able to locate > which tape to load without reading each tape one by one to find the good > one! -> metadata. I wouldn't argue about our system, in part because you're correct, in part because it's not something we can/want to change, as long as it works. Like I said, current nuisance in locating slices doesn't warrant extending it. I still stand by the fact that my argument for slice counters is valid outside of our particular context, otherwise I wouldn't have brought it up. > Doing this inside dar format would lead to add a finite size field to > store slice number, which would limit the number of slices a > backup/archive can have, which I have managed to avoid since day 1 --- > see infinint family classes to handle arbitrarily large integers, a post > Y2K symptom ;). I understand your point that VLQ, i.e. arbitrarily large integer to store slice number is inefficient. I understand that you don't want to put a hard limit on the slice count either. But I don't understand why you argue against a rotating slice counter. You have said earlier that even with the UUID, there is some chance of getting the same value: > There is little chance (not to say no chance at all), that two > different backups have the same internal name. So UUIDs do not guarantee 100% uniqueness (and take way more space than a counter I suggest). Why do you oppose a fixed-width rotating counter so much then? >>> what length (in byte, for example) would you give to this >>> "simple current slice index"? >> >> Practically, 2 bytes and allow it to overflow. I don't have >> statistics > > that's precisely the point: even the following references do not > reflect the variety of use cases dar currently addresses: > > http://dar.linux.free.fr/doc/References.html I agree that we lack exhaustive statistics and there may be exotic cases. But dar has settled on some practical defaults. Slice counter is the same kind of non-definite measure as is the archive label (UUID). UUIDs can duplicate; there may be more slices than 2^16 (2^24, 2^32, ...). Does this fact change their utility? >> Why would it add UUIDs into every slice header if you insist on >> storage integrity? > > to avoid users (as I am) to make a mistake of mixing slices of different > archives without being notified. You oppose me in saying that one cannot reorder slices (even by accident) and that file name is paramount, but you still state that one can mix slices from different archives? Doesn't it require a comparable level of mistake? To accidentally add a slice from another archive, one would need to rename it specifically, right? And not just change <basename> but also <slice number>, otherwise there'd be a noticeable gap or often an overwrite prompt? Even if two directories have slices with the same <basename> and you move a slice from one to another, <slice number>s would still clash. arc.1.dar arc.2.dar arc.3.dar <- real last slice arc.23.dar <- added-in slice from another archive Also, it's usually obvious when something like that happens because last slice is smaller than others so you'd see 2 last slices of different sizes than the rest - unlike when you reorder slices. From this standpoint, to my argument: why one cannot accidentally remove one of the slices (first or last is easiest) or mess up with renaming them (if we're talking about renaming) or duplicating them so that their <slice number> change but <basename> doesn't? Or, if it's directories with slices, fail to move first/last slices to new location? >> Why store UUID then? <basename> is expected to be unique, right? Why >> not use that, trust that the user will make sure to name all slices >> correctly (if you require that user maintains <slice number> himself, >> extending this requirement to <basename>.<slice number> is all the >> more logical) and forego any checks for slice correspondence? > > This is not because you have a seat belt in your car that your are free > and safe to bump into any tree. > > This feature is my seat belt to reduce my time spent in support request > and to ease fault isolation: If users are notified when they mix slices > of different backups/archives, the do not need to and usually don't come > to me. Though, this may not catch all type of human error as a seat belt > cannot prevent all injuries, better having it than nothing. Excellent point. Why doesn't it apply to slice numbers? > I would reverse the question: why limiting the number of slices and > thus the overall backup size (because slice size depends on the > underlying filesystem ability to support them) when you can avoid > it? Like I have repeatedly stated, the counter may be rotating. We are not forced to make it bulletproof just like archive label isn't bulletproof. >>> why a "fixed length numeric field" would be enough in the "2^32 >>> territory" while a "random number" (which size is still >>> undefined) would not? >> >> I assumed that both fields would be of the same size because a >> random number field that's wider than a fixed number field looks >> pointless to me - it takes more space, it doesn't guarantee >> uniqueness even within its range, it doesn't allow comparing >> positions, etc. What's the advantage of such a random field? > > the advantage in that particular case is the fixes width size, which > here is interesting to know at which time, while writing down a > slice, you stop adding data to start adding this random field. I understand your point about variable-length fields. But my question was how a fixed-width random field is better than a fixed-width rotating integer counter. Both do not guarantee uniqueness but the latter has advantages over the former. >> I even mentioned that it might come in handy for other cases (broken >> FAT with unavailable file names). Again, it seems very in-line with >> the existing tape marks and sequential reading mode. > > If you need a robust filesystem or face to filesystem corruptions, this > is a filesystem feature that is needed (RAID, journals, ...) but for > those as I am that do not want so robust filesystem, we do backups ;). You made me understand that tape marks' and -0's potential for recovery is an (undesired) side effect. But why does dar have the lax mode then? Its only purpose is to fight FS corruptions and lack of backups. > Adding a slice number inside each slice would not help much: if a file > (slice) is retrieved in different blocks/files, due to filesystem > corruption, it will not tell you how to stick them back together, > assuming no part is missing, which is improbable in that context. From my experience, and I have tended to a recovery of a dozen or so faulty drives including a SSDs using Windows and Linux file systems, the recovery process of course cannot guess block allocation but usually FAT isn't obliterated entirely - recovery often still produces consistent result for some files. Yet, even if you managed to get a consistent file, you may not manage to get its file name. Moreover, one of the popular reasons for recovery is when a directory was accidentally deleted or a FS was formatted. In this case file entries don't (immediately) disappear and it's possible to fully recover deleted files, but again, their file names may be mangled, fully or partially. I'm sure you're aware of the renowned FOUND.000 directory and how it's populated. https://en.wikipedia.org/wiki/Found.000#cite_ref-14 Many times, I have personally seen entirely consistent files recovered in that directory, but without file names. It still happens with USB sticks, and I can easily imagine dar slices residing on one (it may be irresponsible on the user's behalf but dar already has mechanisms to address irresponsibility, as we've hopefully established). |
|
From: Denis C. <dar...@fr...> - 2026-06-19 15:56:59
|
Le 18/06/2026 à 22:45, mannino a écrit : > Hi Denis, Hi Mannino, > > We have powered down both systems and ran memtest86+ for one full pass. > No RAM errors were detected. Good thing. Thanks for having taken the time to test it. This means also that the hash (--hash) generated by dar should be significant, as they are computed in memory, thus you can further know whether storage has brought some corruption or not. > > Just before the shutdown, one of those systems has created a full dar > archive with the same kind of corruption (direct mode doesn't work, -t > -0 reports 5 CRC errors). After powering it on and manually running full > dar backup, a valid archive was produced. this is quite weird... we get close to the invocation of the cosmic particle that have changed something somewhere in some memory resident code... (I'm almost kidding!) > > Since we seem to be out of ideas, I will have to keep an eye on the > backups coming from these two systems and let you know if it happens > again. I have enabled --hash and we now know that the RAM is fine so > we'd be a step forward if it does happen. exactly > > NB: typo on dar(1): "These hash files can be processes by md5sum" - s/ > processes/processed/ Oh! I'll will fix that shortly, thanks you for the feedback. [...] > >>>>> 2. Byte sizes of non/last slices would allow basic integrity check but >>>>> (1) `file` doesn't show them; (2) dar doesn't show them in -l -0 mode >>>>> (so unable to extract from invalid file); (3) dar -l -q works on >>>>> isolated catalogue but only reports non-last slice size - given our >>>>> problem here is likely last slice being truncated, it doesn't really >>>>> help. >> >> dar has no need to show this information to work normally, either this >> is the expected slice or a warning shows to the users for they take the >> corrective action. > > It already shows a lot of info (including sizes) in -l -q mode. Why not > make it available in sequential read mode? because this information is not stored inside the archive data, but found in the metadata (here the filenames) with in addition the side effect of code modularity: Follows the modularity explained in a long answer, skip over it if you don't want to known the libdar internals: at the bottom floor takes place either: - a sar object (segment and reassembly) which manages slices and exposes to the upper level a virtual large file - or a trivial_sar object, which mimics the sar behavior when the data is read from a pipe - a zapette which is a remote control of a sar or trivial_sar hosted in dar_slave program. At the next floors, you will find some or all possible options: - an encryption layer (of a cache layer if no encryption is used) - escape layer which adds/read tape mark needed to interlace file's data and their metadata - compression layer on top. This stack of objects, reading from each others, is provided to an archive object (which realizes the floor 3 to 6 depending on the existence of layers below it) and this object just have access to: - the unciphered, desinterlaced, decompressed data in the form of a virtual single big file containing all data and EA of all saved files. - and as an option (if tape marks are present in the archive) an access to the escape layer to seek forward (floor 2) for the next mark of a given type. In direct mode the listing process only involves reading the catalogue stored in memory, which has been loaded by opening the last slice at which time libdar can get the last slice number and its size, which none are stored in the archive, because storing them might change them if for example this addition would create a new slice. In sequential read mode, at the top of the stack the archive avoids reading any filesystem metadata, it does not seek() forward or backward in this virtual big file, but just read(). Eventually it uses the handle provided by the escape layer to read forward up to the next tape mark of a given type (to fetch the next file metadata: filename, ownership, permission, dates...) stored in-band. But as the stack may be sitting on a named or even anonymous pipe for which there is not such filesystem metadata as the filename, the code avoids seeking() at a given position and works the same whatever the underneath stack content is (sliced, unsliced, pipe, remote control of dar_slave). Moreover, both approaches share the same filtering code (its one level above the archive) and here the below layers are completely abstracted. It only has access to the catalogue (or the escape_catalogue object in sequential-read mode) read the next entry, consider the filter used (based on filename, path, EA...) and eventually do something with it (listing or fetching the data for testing, restoration,...) before looping with the next archive entry. What is abstracted at this layer is the way the saved file metatada is fetched: - in direct mode this is already loaded in memory in the catalogue object, this is just a memory read in this data-structure. - in sequential-read mode, the escape_catalogue when asked for the next entry, drives the escape_layer to read() up to the next mark, fetches reads the metadata and builds a catalogue in memory during this process. (escape_catalogue inherits from class escape). In sequential read mode when listing the details of an archive (-l -q), the offset and size of the last slice are not available because the below layer may be just a pipe nor a file and this code must works the same on pipes, single sliced, multi-sliced archive or remote control for dar_slave. So yes, this is a long explanation to say I will not break this abstraction making a hole in all the floors to expose something that has not to be know at this a higher level. Having abstraction between parts of code ease maintenance, evolution and reduce the risk of bugs. So I'm sorry I'll not break that. Here I will rather hide behind the "design" principle and the fact a design is a choice, a policy. [...] >>> >>> Okay, the ability to view the details of slice-layout in -0 mode would >>> be welcome. I'd say any header info available to dar should be made >>> visible on equal terms. >> >> I guess when you drive a car, you don't seen all the internal sensor >> values (there is a lot of them in today's car), but over the classical >> few indicators (oil temperature, speed, motor RPM,...) this is only when >> something wrong happens that these internal sensors trigger an alarm on >> your car dashboard in human understandable form, no? > > This is why I'm always adding "in -0 mode". This mode is exactly for > "when something wrong happens", right? No. The sequential-read mode is mainly to cope with devices that do not provide direct access mode to a given offset of a file. It was absolutely not part of the initial design but could be added --- thanks to libdar internal abstractions and modularity --- to address mainly tape support and as a side use case the reading through any type of piped data. Today it is discouraged to remove tape marks while this would speed up the backup and reading processes, because tape marks (with the metadata that follows) can also be used to rebuild the internal catalogue of a truncated archive (-y option), it provides saved files metadata redundancy within a given archive. However you should consider the sequential-read mode as restricted mode of operation (due to the tape/pipe context that is the target here). The direct (and original) mode of operation is much more efficient as it directly reads and only reads the needed data for restoration, testing, listing, diffing,... especially when using filtering mechanisms. > >> For dar this is the same, first it would need a lot of work to expose >> everything and second it would lead to an unreadable output, dar is >> already too much talkative to my standpoint and I guess to the common >> opinion. > > Verbose header info output is meant for specific modes (-l -q), not > regular modus operandi. I've added the following in the TODO list (at sourceforge) : expose the internal_name and data_name a given archive in that -l -q mode (or -l -v). This will be a per archive information and not per slice, as dar does not handle slices individually. > >> Another mistake you apparently do is to consider dar would manipulate a >> slice individually. Dar manipulate at once a set of slice, *by design*. >> Using the sequential-read mode to feed a arbitrary slice in the hope to >> get some information of it is just misuse of dar. > > I never said I wanted to feed random slices to dar. I only said that > "the ability to view the details of slice-layout in -0 mode would > be welcome". Currently, -l -q works but -l -q -0 doesn't, even if I feed > it the first slice. OK, I misunderstood what you expected to do, sorry. [...] >> >> If by "bucket" you mean those of an S3 storage or Azure Blob Storage, >> why not using the metadata you can assign to these objects to record the >> slice information??? >> >> This would pretty make sense, no? > > We don't use S3, we use a homebrew system. I only used the term "bucket" > as our system somewhat resembles an object storage. I think your homebrew system is missing metadata support. Without it, you are stuck to the them same approach as tapes and all the restrictions it has. In fact, your homebrew system is probably even more restrictive than a tape system, as users usually add a sticker on their tapes when manually handled or the robot is hopefully able to locate which tape to load without reading each tape one by one to find the good one! -> metadata. Now, I'm sorry to say so but, if you need to keep this metadata associated to the slice data (metadata dar relies on but which you suppressed when you removed the slice filenames) that's your duty to have your howebrew storage system carrying this need, either by: - wrapping the slices within a file format that contains the slice number followed by the slice content (thought this would be inefficient as it would require to load and read a slice to know whether this is the good one), - or add support for metadata beside the file/slice data to store this information (slice number). Doing this inside dar format would lead to add a finite size field to store slice number, which would limit the number of slices a backup/archive can have, which I have managed to avoid since day 1 --- see infinint family classes to handle arbitrarily large integers, a post Y2K symptom ;). [...] >> So to face some data corruption, you ask all applications to add the >> necessary redundant information that could be lost by the storage >> solution? > > I merely propose to further recovery mode (-0) already exiting in dar. > Why would dar add tape marks if it trusted the storage? Not a lack of trust toward the storage, but a lack of feature: Tape marks are used to identify the next file just by using the read() system call, this is needed when the underlying media has not support to seek() to a given offset (pipes) or has so poor performance doing so that it is better using read (tape drives). > Why would it add > UUIDs into every slice header if you insist on storage integrity? to avoid users (as I am) to make a mistake of mixing slices of different archives without being notified. > Why > implement -0 in the first place with this mindset? as written above -0 mode was added later to address tape and pipe support. > >> I think you have to take dar for what it is: a tool that expects slices >> names of that format <basename>.<slice number>.dar which is however >> flexible enough for those slices be fed through pipes or other >> mechanisms, like file regeneration based on blobs or other opaque data >> structure as you described. But this is on the storage side to handle >> failures and provide the necessary environment to adapt to the tool >> which are expected to be used... not the inverse. > > Why store UUID then? <basename> is expected to be unique, right? Why not > use that, trust that the user will make sure to name all slices > correctly (if you require that user maintains <slice number> himself, > extending this requirement to <basename>.<slice number> is all the more > logical) and forego any checks for slice correspondence? This is not because you have a seat belt in your car that your are free and safe to bump into any tree. This feature is my seat belt to reduce my time spent in support request and to ease fault isolation: If users are notified when they mix slices of different backups/archives, the do not need to and usually don't come to me. Though, this may not catch all type of human error as a seat belt cannot prevent all injuries, better having it than nothing. > >> Moreover reading blobs metadata is very light, while instead fetching >> the many megabytes of a object/blob/slice to just pass it to a tool >> (file command or dar) that would tells you this is not the good slice is >> un-optimal at all and quite costly ($), no? > > It is, but we avoid that by examining date of admission since there's a > predictable time gap between incoming slices and archives. I use `file` > after fetching the data only when investigating issues. Sorry to say so again, but that's your choice and design (and policy), I respect it and I will no more discuss about the metadata support I think your homebrew storage is lacking. On my side too I have my design, choice and policy with dar/libdar: I don't want to support user specific/homebrew storage systems. [...] >> >> same remark as above: the bucket/object/blob metadata is the logical >> place of the slice information you remove from filesystem metadata (= >> filename). >> >> You have to find a way to translate from filesystem metadata to >> S3/bucket/objet/blob metadata and vice versa, this is not the purpose of >> to do that very context specific task. > > You seem to be constantly misinterpreting my words in spite of me > explicitly stating that all my suggestions are concerning recovery mode, > not normal operation. it was not intentional, sorry > We have been using our (arguably weird) storage > system along with dar for over a decade and it works just fine - without > dar's built-in slice counters, extra UUIDs, etc. But our current problem > highlighted what I believe is a shortcoming on dar's diagnostic/recovery > side and I simply propose some improvements (very modest really). As mentioned above, I have added a feature request in the todo list to expose the internal_name and data_names of an archive when using -l -q (or -l -v), > I even > mentioned that it might come in handy for other cases (broken FAT with > unavailable file names). Again, it seems very in-line with the existing > tape marks and sequential reading mode. If you need a robust filesystem or face to filesystem corruptions, this is a filesystem feature that is needed (RAID, journals, ...) but for those as I am that do not want so robust filesystem, we do backups ;). Adding a slice number inside each slice would not help much: if a file (slice) is retrieved in different blocks/files, due to filesystem corruption, it will not tell you how to stick them back together, assuming no part is missing, which is improbable in that context. > >>>> why only 4 bytes? This would limit the number of slices... even if >>>> today >>>> you think it is large enough, this is the same way of thinking that led >>>> to the year 2000 bug, the problem of year 2038, the upper and high >>>> memory in DOS some decades ago, the 2 GB files boundary (and need of >>>> large file support)... and so forth. >>> >>> I understand your point. But even if we take 1 KiB per slice, 2^32 >>> slices would add up to 4.4 TiB of data... and more "realistic" 1 MiB >>> per slice would be 4.5 PiB... but if that is a concern, you could use >>> 64 bits, that would be (let me look that up on Wikipedia) around 2 >>> exabytes of data if each slice only held 1 byte of payload... I'm >>> afraid we'd need more than quantum computers to cross that boundary... >> >> an so now, the smallest slice which is a few tens of bytes would becomes >> a few kilobytes... what a waste... > > What's the practical reason for having over 4 billions of slices, each a > few tens of bytes? Besides, the counter may be rotating, it's only a > hint anyway. this is what they say when storing years on 2 digits last century, or sotring dates as a number of seconds since 1969 ended in a 32 bit fields (year 2038 problem, we are getting to it...), and so forth. I would reverse the question: why limiting the number of slices and thus the overall backup size (because slice size depends on the underlying filesystem ability to support them) when you can avoid it? > >>>> A better approach could be to add a random fixed size field at end of >>>> slice and have this same random value at the beginning of the next. But >>>> I don't see the use case as mentioned earlier about number of slices in >>>> filenames. >>> >>> If we're in the territory where 2^32 slices may not be enough, our >>> random numbers may not be random enough too, especially within >>> millions of slices... If we were choosing, I'd just use a fixed length >>> numeric field for the sake of simplicity. VLQ if you must, although >>> I'd grade that as an overkill... >> >> why a "fixed length numeric field" would be enough in the "2^32 >> territory" while a "random number" (which size is still undefined) would >> not? > > I assumed that both fields would be of the same size because a random > number field that's wider than a fixed number field looks pointless to > me - it takes more space, it doesn't guarantee uniqueness even within > its range, it doesn't allow comparing positions, etc. What's the > advantage of such a random field? the advantage in that particular case is the fixes width size, which here is interesting to know at which time, while writing down a slice, you stop adding data to start adding this random field. It is also necessary at reading time to translate the location of a saved file data, stored as an offset in the virtual big file I explained above, into a slice number and offset in that slice. Having variable width field to store slices number would make it impossible without reading all slices header and trailers... and storing slice sizes in a fixed width field would look like one shooting at his foot. > >> what length (in byte, for example) would you give to this "simple >> current slice index"? > > Practically, 2 bytes and allow it to overflow. I don't have statistics that's precisely the point: even the following references do not reflect the variety of use cases dar currently addresses: http://dar.linux.free.fr/doc/References.html I will this not limit this to address a user specific context. > but I don't think 65k+ slices per archive are a norm, and even if this > practice exists somewhere, a rotating 16-bit counter would make it 2^16 > times less likely that an incorrect ordering of slices would go > undetected (and 2^16 times easier to determine the number of any given > slice) compared to when there's no such counter at all. On dar's side, > the check would simply go from 0 to 2^16, then overflow and repeat. > Regards, Denis |
|
From: mannino <ma...@of...> - 2026-06-18 20:46:28
|
Hi Denis, We have powered down both systems and ran memtest86+ for one full pass. No RAM errors were detected. Just before the shutdown, one of those systems has created a full dar archive with the same kind of corruption (direct mode doesn't work, -t -0 reports 5 CRC errors). After powering it on and manually running full dar backup, a valid archive was produced. Since we seem to be out of ideas, I will have to keep an eye on the backups coming from these two systems and let you know if it happens again. I have enabled --hash and we now know that the RAM is fine so we'd be a step forward if it does happen. NB: typo on dar(1): "These hash files can be processes by md5sum" - s/processes/processed/ > it is partially the case in 2.8.x and is now fully the case in 2.9.x > (this is the feature I was working on these last days). So now an > isolated catalogue has all the needed information to operate the backup > of reference and does not more need the first or last slice to know the > slice-layout, archive format, compression algo, ciphering salt and other > parameters of that sort. Excellent addition. >>>> 2. Byte sizes of non/last slices would allow basic integrity check but >>>> (1) `file` doesn't show them; (2) dar doesn't show them in -l -0 mode >>>> (so unable to extract from invalid file); (3) dar -l -q works on >>>> isolated catalogue but only reports non-last slice size - given our >>>> problem here is likely last slice being truncated, it doesn't really >>>> help. > > dar has no need to show this information to work normally, either this > is the expected slice or a warning shows to the users for they take the > corrective action. It already shows a lot of info (including sizes) in -l -q mode. Why not make it available in sequential read mode? >>> The sar layer (which manages slices) >>> processes a stream of bytes from upper layers (ciphering, compression, >>> filtering, and so forth) which is generated on the fly from the >>> filesystem to backup, the sar takes this flow and creates slices from it >>> up to the time the flow dries up. >>> >>> What is stored at the beginning of *each slice* is the so called >>> slice-layout: >>> - first slice max size, >>> - other slices max size, >>> - slice header size >>> - and slice format (2 formats so far). >>> >>> Only this could be shown in sequential-read mode, not what is displayed >>> today in direct mode which adds to this the overall size of the archive >>> and the size of the last slice (easy to get because in that mode this is >>> the last slice that get read first). >> >> Okay, the ability to view the details of slice-layout in -0 mode would >> be welcome. I'd say any header info available to dar should be made >> visible on equal terms. > > I guess when you drive a car, you don't seen all the internal sensor > values (there is a lot of them in today's car), but over the classical > few indicators (oil temperature, speed, motor RPM,...) this is only when > something wrong happens that these internal sensors trigger an alarm on > your car dashboard in human understandable form, no? This is why I'm always adding "in -0 mode". This mode is exactly for "when something wrong happens", right? > For dar this is the same, first it would need a lot of work to expose > everything and second it would lead to an unreadable output, dar is > already too much talkative to my standpoint and I guess to the common > opinion. Verbose header info output is meant for specific modes (-l -q), not regular modus operandi. > Another mistake you apparently do is to consider dar would manipulate a > slice individually. Dar manipulate at once a set of slice, *by design*. > Using the sequential-read mode to feed a arbitrary slice in the hope to > get some information of it is just misuse of dar. I never said I wanted to feed random slices to dar. I only said that "the ability to view the details of slice-layout in -0 mode would be welcome". Currently, -l -q works but -l -q -0 doesn't, even if I feed it the first slice. >>>> 3. Slice number is what I truly miss. According to my tests, messing >>>> with non-last slices (removing, duplicating, reordering) is usually >>>> detected by dar as different types of data inconsistency so even if >>>> you have all slices perfectly good and well, just in the wrong order, >>>> you would think the backup is botched, and those errors would only be >>>> misleading. >>> >>> the slice number is expected to be in the filename and it is not >>> expected that users would change/renumber it... >> >> Let me explain our storage strategy briefly. Backup files are held off >> site for obvious reasons; the storage space is a set of buckets (could >> be one or more per server but they're static) storing opaque blobs >> whose only characteristics are acceptance date and data itself. Date >> is stored for FIFO-like purging. Blobs can be moved or duplicated >> elsewhere behind the scenes. >> >> I understand it may not be the most elegant or efficient solution but >> that's what we have now. Allowing free form metadata like file names >> would complicate blob management. For backup purposes, this is not a >> problem because each backup location is tied to a particular bucket >> where archives and slices go in chronological order so by inspecting >> date and data size one can clearly divide entries into individual >> archives even if there are no file names. We store additional data in >> dar's user comment field, but it's obviously archive-wise. > > If by "bucket" you mean those of an S3 storage or Azure Blob Storage, > why not using the metadata you can assign to these objects to record the > slice information??? > > This would pretty make sense, no? We don't use S3, we use a homebrew system. I only used the term "bucket" as our system somewhat resembles an object storage. >>> This is something unclear to me: you have this slice number information >>> in the slice filename... isn't it the easiest way to know what slice >>> number is a slice? And how much slice do one have? And also what is the >>> latest slice? >> >> My point is that in some scenarios file names are unavailable. Even if >> not for our storage system, file names may be missing as a result of >> data recovery from a faulty drive, for example. Sure, that should not >> normally happen, and if you use par2 then it would fix the names for >> you as well, but we're talking about last-chance recovery here, and >> that's why dar has the -0 mode, so why not take it a step further with >> one or two simple new fields in the header? > > So to face some data corruption, you ask all applications to add the > necessary redundant information that could be lost by the storage solution? I merely propose to further recovery mode (-0) already exiting in dar. Why would dar add tape marks if it trusted the storage? Why would it add UUIDs into every slice header if you insist on storage integrity? Why implement -0 in the first place with this mindset? > I think you have to take dar for what it is: a tool that expects slices > names of that format <basename>.<slice number>.dar which is however > flexible enough for those slices be fed through pipes or other > mechanisms, like file regeneration based on blobs or other opaque data > structure as you described. But this is on the storage side to handle > failures and provide the necessary environment to adapt to the tool > which are expected to be used... not the inverse. Why store UUID then? <basename> is expected to be unique, right? Why not use that, trust that the user will make sure to name all slices correctly (if you require that user maintains <slice number> himself, extending this requirement to <basename>.<slice number> is all the more logical) and forego any checks for slice correspondence? > Moreover reading blobs metadata is very light, while instead fetching > the many megabytes of a object/blob/slice to just pass it to a tool > (file command or dar) that would tells you this is not the good slice is > un-optimal at all and quite costly ($), no? It is, but we avoid that by examining date of admission since there's a predictable time gap between incoming slices and archives. I use `file` after fetching the data only when investigating issues. >>> What I don't understand is how you could >>> accidentally miss the first slice, I mean how do you let dar accessing >>> the slices as dar, even in sequential-read mode open a slice after a >>> given filename and if it is not present it complains and waits. >> >> To deploy or even inspect a backup, its blobs should be pulled from >> storage somewhere dar can run, then renamed or symlinked to form >> arc.N.dar. Human factor makes such mistakes possible. If you forgot >> one or more leading slices yet they're named arc.1.dar, arc.2.dar, >> etc. - dar won't "complain and wait", it'd proceed straight to >> throwing errors at you. > > same remark as above: the bucket/object/blob metadata is the logical > place of the slice information you remove from filesystem metadata (= > filename). > > You have to find a way to translate from filesystem metadata to > S3/bucket/objet/blob metadata and vice versa, this is not the purpose of > to do that very context specific task. You seem to be constantly misinterpreting my words in spite of me explicitly stating that all my suggestions are concerning recovery mode, not normal operation. We have been using our (arguably weird) storage system along with dar for over a decade and it works just fine - without dar's built-in slice counters, extra UUIDs, etc. But our current problem highlighted what I believe is a shortcoming on dar's diagnostic/recovery side and I simply propose some improvements (very modest really). I even mentioned that it might come in handy for other cases (broken FAT with unavailable file names). Again, it seems very in-line with the existing tape marks and sequential reading mode. >>> why only 4 bytes? This would limit the number of slices... even if today >>> you think it is large enough, this is the same way of thinking that led >>> to the year 2000 bug, the problem of year 2038, the upper and high >>> memory in DOS some decades ago, the 2 GB files boundary (and need of >>> large file support)... and so forth. >> >> I understand your point. But even if we take 1 KiB per slice, 2^32 >> slices would add up to 4.4 TiB of data... and more "realistic" 1 MiB >> per slice would be 4.5 PiB... but if that is a concern, you could use >> 64 bits, that would be (let me look that up on Wikipedia) around 2 >> exabytes of data if each slice only held 1 byte of payload... I'm >> afraid we'd need more than quantum computers to cross that boundary... > > an so now, the smallest slice which is a few tens of bytes would becomes > a few kilobytes... what a waste... What's the practical reason for having over 4 billions of slices, each a few tens of bytes? Besides, the counter may be rotating, it's only a hint anyway. >>> A better approach could be to add a random fixed size field at end of >>> slice and have this same random value at the beginning of the next. But >>> I don't see the use case as mentioned earlier about number of slices in >>> filenames. >> >> If we're in the territory where 2^32 slices may not be enough, our >> random numbers may not be random enough too, especially within >> millions of slices... If we were choosing, I'd just use a fixed length >> numeric field for the sake of simplicity. VLQ if you must, although >> I'd grade that as an overkill... > > why a "fixed length numeric field" would be enough in the "2^32 > territory" while a "random number" (which size is still undefined) would > not? I assumed that both fields would be of the same size because a random number field that's wider than a fixed number field looks pointless to me - it takes more space, it doesn't guarantee uniqueness even within its range, it doesn't allow comparing positions, etc. What's the advantage of such a random field? > what length (in byte, for example) would you give to this "simple > current slice index"? Practically, 2 bytes and allow it to overflow. I don't have statistics but I don't think 65k+ slices per archive are a norm, and even if this practice exists somewhere, a rotating 16-bit counter would make it 2^16 times less likely that an incorrect ordering of slices would go undetected (and 2^16 times easier to determine the number of any given slice) compared to when there's no such counter at all. On dar's side, the check would simply go from 0 to 2^16, then overflow and repeat. |
|
From: Denis C. <dar...@fr...> - 2026-06-14 22:03:36
|
Le 03/06/2026 à 21:49, mannino a écrit : Hi, I have finally found some time to answer the feature related part or your email: >>> 1. Archive UUID, as it's currently implemented, is perfect: (1) it >>> practically guarantees that one cannot mix up slices from different >>> archives; (2) dar itself checks that when testing; (3) can be >>> displayed with `file`: >>> >>> arc.10.dar: dar archive, label "82af71c2 00000000 fffff752" >>> >>> I could only suggest storing UUID of the reference archive in the -@ >>> catalogue too (since -@ gets its own different UUID) and for dar >>> itself to provide some way to view this UUID. In any case, UUID helps >>> one work out the consistency among archives - but there's nothing >>> similar for working out slices for the archive itself, and it's a >>> problem for me. it is partially the case in 2.8.x and is now fully the case in 2.9.x (this is the feature I was working on these last days). So now an isolated catalogue has all the needed information to operate the backup of reference and does not more need the first or last slice to know the slice-layout, archive format, compression algo, ciphering salt and other parameters of that sort. >> >> if you have a look at dar's documentation about the archive format, you >> will see that this is done with more flexibility: >> - internal name : links slices of an archive together >> - data name : links a catalog (or isolated catalog) to the data set to >> use. This prevents one trying to restore an archive with the help of an >> unrelated isolated catalog. >> >> http://dar.linux.free.fr/doc/Notes.html#archive_structure > [...] > Now, if only > dar could provide some way to view this info, preferably in -0 mode for > potentially corrupted slices. this is in the pipe of features: the data_name and internal_name will be shown when listing the detail of a backup. However this information will not tell you whether a slice is corrupted or not. The slice header takes a few tens of bytes at the beginning of each slice (and 1 byte is used at the end of each slice). Corruption taking place in the middle will be reported the same as today (CRC of the data or random other error depending on the place the corruption occurs). Last, dar manipulates a backup (= a set of slices as a whole) it will report this information as fetched from the first slice seen for the listing operation (the first in sequential more, the last in default mode) any other slices involved in the operation will either match both slices attributes or an error will be reported as of today in that same situation. If you want the file command to display this information for any slice individually, you could ask such feature to the maintainer of this command. [...] > >>> 2. Byte sizes of non/last slices would allow basic integrity check but >>> (1) `file` doesn't show them; (2) dar doesn't show them in -l -0 mode >>> (so unable to extract from invalid file); (3) dar -l -q works on >>> isolated catalogue but only reports non-last slice size - given our >>> problem here is likely last slice being truncated, it doesn't really >>> help. dar has no need to show this information to work normally, either this is the expected slice or a warning shows to the users for they take the corrective action. >> >> dar cannot know in advance how much slices will be necessary nor what >> will be the size of the last slice. > > Of course it's streaming so I'm not suggesting storing total number of > slices in each slice. My suggestion here is to store each slice's size - > which is of course known beforehand. (I've written "non/last" but I > shall recede from the "last" part indeed.) All slices except the last one of a backup (= set of slices) have a known size which is described in the slice layout already stored at the beginning of each slice. > >> The sar layer (which manages slices) >> processes a stream of bytes from upper layers (ciphering, compression, >> filtering, and so forth) which is generated on the fly from the >> filesystem to backup, the sar takes this flow and creates slices from it >> up to the time the flow dries up. >> >> What is stored at the beginning of *each slice* is the so called >> slice-layout: >> - first slice max size, >> - other slices max size, >> - slice header size >> - and slice format (2 formats so far). >> >> Only this could be shown in sequential-read mode, not what is displayed >> today in direct mode which adds to this the overall size of the archive >> and the size of the last slice (easy to get because in that mode this is >> the last slice that get read first). > > Okay, the ability to view the details of slice-layout in -0 mode would > be welcome. I'd say any header info available to dar should be made > visible on equal terms. I guess when you drive a car, you don't seen all the internal sensor values (there is a lot of them in today's car), but over the classical few indicators (oil temperature, speed, motor RPM,...) this is only when something wrong happens that these internal sensors trigger an alarm on your car dashboard in human understandable form, no? For dar this is the same, first it would need a lot of work to expose everything and second it would lead to an unreadable output, dar is already too much talkative to my standpoint and I guess to the common opinion. Another mistake you apparently do is to consider dar would manipulate a slice individually. Dar manipulate at once a set of slice, *by design*. Using the sequential-read mode to feed a arbitrary slice in the hope to get some information of it is just misuse of dar. > >>> 3. Slice number is what I truly miss. According to my tests, messing >>> with non-last slices (removing, duplicating, reordering) is usually >>> detected by dar as different types of data inconsistency so even if >>> you have all slices perfectly good and well, just in the wrong order, >>> you would think the backup is botched, and those errors would only be >>> misleading. >> >> the slice number is expected to be in the filename and it is not >> expected that users would change/renumber it... > > Let me explain our storage strategy briefly. Backup files are held off > site for obvious reasons; the storage space is a set of buckets (could > be one or more per server but they're static) storing opaque blobs whose > only characteristics are acceptance date and data itself. Date is stored > for FIFO-like purging. Blobs can be moved or duplicated elsewhere behind > the scenes. > > I understand it may not be the most elegant or efficient solution but > that's what we have now. Allowing free form metadata like file names > would complicate blob management. For backup purposes, this is not a > problem because each backup location is tied to a particular bucket > where archives and slices go in chronological order so by inspecting > date and data size one can clearly divide entries into individual > archives even if there are no file names. We store additional data in > dar's user comment field, but it's obviously archive-wise. If by "bucket" you mean those of an S3 storage or Azure Blob Storage, why not using the metadata you can assign to these objects to record the slice information??? This would pretty make sense, no? > > So essentially our backup directory looks like this: > > 402 | Sun Feb 22 11:43:58 AM UTC 2026 | 167511621 bytes > 403 | Tue Mar 3 06:56:46 PM UTC 2026 | 10613869728 bytes > 404 | Tue Mar 3 07:13:29 PM UTC 2026 | 10613869728 bytes > 405 | Tue Mar 3 07:36:56 PM UTC 2026 | 635571924 bytes > >> This is something unclear to me: you have this slice number information >> in the slice filename... isn't it the easiest way to know what slice >> number is a slice? And how much slice do one have? And also what is the >> latest slice? > > My point is that in some scenarios file names are unavailable. Even if > not for our storage system, file names may be missing as a result of > data recovery from a faulty drive, for example. Sure, that should not > normally happen, and if you use par2 then it would fix the names for you > as well, but we're talking about last-chance recovery here, and that's > why dar has the -0 mode, so why not take it a step further with one or > two simple new fields in the header? So to face some data corruption, you ask all applications to add the necessary redundant information that could be lost by the storage solution? I think you have to take dar for what it is: a tool that expects slices names of that format <basename>.<slice number>.dar which is however flexible enough for those slices be fed through pipes or other mechanisms, like file regeneration based on blobs or other opaque data structure as you described. But this is on the storage side to handle failures and provide the necessary environment to adapt to the tool which are expected to be used... not the inverse. Moreover reading blobs metadata is very light, while instead fetching the many megabytes of a object/blob/slice to just pass it to a tool (file command or dar) that would tells you this is not the good slice is un-optimal at all and quite costly ($), no? > >> What I don't understand is how you could >> accidentally miss the first slice, I mean how do you let dar accessing >> the slices as dar, even in sequential-read mode open a slice after a >> given filename and if it is not present it complains and waits. > > To deploy or even inspect a backup, its blobs should be pulled from > storage somewhere dar can run, then renamed or symlinked to form > arc.N.dar. Human factor makes such mistakes possible. If you forgot one > or more leading slices yet they're named arc.1.dar, arc.2.dar, etc. - > dar won't "complain and wait", it'd proceed straight to throwing errors > at you. same remark as above: the bucket/object/blob metadata is the logical place of the slice information you remove from filesystem metadata (= filename). You have to find a way to translate from filesystem metadata to S3/bucket/objet/blob metadata and vice versa, this is not the purpose of to do that very context specific task. > >>> Also, out of curiosity, I tried copying one slice twice and got a ton >>> of CRC errors. >> > > With the help of 4 extra bytes (= slice number) in each slice's >> header, >>> we would be able to identify the problem precisely. >> >> why only 4 bytes? This would limit the number of slices... even if today >> you think it is large enough, this is the same way of thinking that led >> to the year 2000 bug, the problem of year 2038, the upper and high >> memory in DOS some decades ago, the 2 GB files boundary (and need of >> large file support)... and so forth. > > I understand your point. But even if we take 1 KiB per slice, 2^32 > slices would add up to 4.4 TiB of data... and more "realistic" 1 MiB per > slice would be 4.5 PiB... but if that is a concern, you could use 64 > bits, that would be (let me look that up on Wikipedia) around 2 exabytes > of data if each slice only held 1 byte of payload... I'm afraid we'd > need more than quantum computers to cross that boundary... an so now, the smallest slice which is a few tens of bytes would becomes a few kilobytes... what a waste... > >> A better approach could be to add a random fixed size field at end of >> slice and have this same random value at the beginning of the next. But >> I don't see the use case as mentioned earlier about number of slices in >> filenames. > > If we're in the territory where 2^32 slices may not be enough, our > random numbers may not be random enough too, especially within millions > of slices... If we were choosing, I'd just use a fixed length numeric > field for the sake of simplicity. VLQ if you must, although I'd grade > that as an overkill... why a "fixed length numeric field" would be enough in the "2^32 territory" while a "random number" (which size is still undefined) would not? I don't argue for this feature or not just don't understand your reasoning... > >>> Not only will it tell us slice order, it will also tell us the total >>> number of slices (since identifying last slice is trivial). >> >> not to my standpoint, as all slices may not be visible at a given time >> in particular when using removal media or remote storage. > > Again, I'm not advocating storing *total* number of slices anywhere. My > point is that, with a simple "current slice index" field embedded in > each slice-layout, it's possible to eliminate all slice order-related > errors and usually even derive the total number of slices by taking last > slice's counter (given that last slice is smaller than others, detecting > it empirically is trivial). Sparing several bytes in each slice is a > small price to pay for this benefit, IMHO. what length (in byte, for example) would you give to this "simple current slice index"? > >>> Storing total number of slices in the catalogue would take it even >>> further, >> >> this would break the ability to re-slice (dar_xform) an archive, once >> the isolated catalogue would have been created. > > Okay, no worries, I take this suggestion back. > > Thank you for still being on board with me on this! > Cheers, Denis |
|
From: Denis C. <dar...@fr...> - 2026-06-06 19:45:19
|
Le 06/06/2026 à 21:32, mannino a écrit : > >> Don't take any offense from my last and short message. > > I didn't. > >> The time scale for support (hours, days) is not the same as the one for >> feature addition (weeks, months, years). > > I understand that. > >> I cannot thus add a feature to just understand what's wrong in a backup >> process > > Adding the counter will not help in figuring already existing backups. > That wasn't the point. The point was to suggest/discuss adding it > potentially. Not this week or year even. Eventually. Is this a > possibility? I tried to indicate its utility outside of my particular > setup but you just discarded that half of my message and I'm not sure > why. If you've already decided against it then I won't take offense at > the straight answer. Otherwise, we can start a new thread for this if > you prefer. > i have not decided anything yet, just not the same problem/time scale. Please fill a feature request here https://sourceforge.net/p/dar/feature-requests/ I will consider when the time has come, no worries. |
|
From: mannino <ma...@of...> - 2026-06-06 19:32:58
|
> Don't take any offense from my last and short message. I didn't. > The time scale for support (hours, days) is not the same as the one for > feature addition (weeks, months, years). I understand that. > I cannot thus add a feature to just understand what's wrong in a backup > process Adding the counter will not help in figuring already existing backups. That wasn't the point. The point was to suggest/discuss adding it potentially. Not this week or year even. Eventually. Is this a possibility? I tried to indicate its utility outside of my particular setup but you just discarded that half of my message and I'm not sure why. If you've already decided against it then I won't take offense at the straight answer. Otherwise, we can start a new thread for this if you prefer. |
|
From: Denis C. <dar...@fr...> - 2026-06-06 13:50:57
|
Le 05/06/2026 à 23:30, mannino a écrit : > I understand. Thank you for your responses. I shall look into --hash, > it's very good that it digests on the fly. I shall also get RAM of one > or both our servers checked whenever possible. > > From your standpoint, are both of my archives broken in a similar way > (i.e. including the archive from my last message)? > > As for my suggestion of storing slice counter in the header, you're not > considering it after all? That reasoning should stand regardless of my > particular problem. > Don't take any offense from my last and short message. The time scale for support (hours, days) is not the same as the one for feature addition (weeks, months, years). I cannot thus add a feature to just understand what's wrong in a backup process, in particular when we reached a step during the investigations, where the most probable cause (even if it sounds improbable) seems to be outside dar/libdar, until this probable cause is invalidated or not by real test (by facts, not by assumptions). On the other hand, adding a feature needs planning, consideration of already existing features, architecture review, then code review, before defining how it will be implemented... the comes the coding time. Note also that I have limited time available for dar/libdar and spending it to help for support requests is less time available for new features. Last, dar release process is long as it includes a lot of testing and adds some code reviews... for me this is the guaranty to not be overwhelmed by support requests and by user dissatisfaction (in particular of myself as being also a user of dar!) The testing passed include all the features you have been using here, I'm confident enough that seen the amplitude of the corruption you have, I (an the many other users) would have also met it if it was just a bug dar/libdar... but the we will see. Now, in you context, I explain what would be the best next step to follow according to my understanding of dar code and of the context you have... that's all. Cheers, Denis |
|
From: mannino <ma...@of...> - 2026-06-05 21:31:19
|
I understand. Thank you for your responses. I shall look into --hash, it's very good that it digests on the fly. I shall also get RAM of one or both our servers checked whenever possible. From your standpoint, are both of my archives broken in a similar way (i.e. including the archive from my last message)? As for my suggestion of storing slice counter in the header, you're not considering it after all? That reasoning should stand regardless of my particular problem. |
|
From: Denis C. <dar...@fr...> - 2026-06-05 20:54:24
|
Le 03/06/2026 à 21:49, mannino a écrit : > [...] > > Meanwhile, you suggested we rule out disk errors, transfer errors and > wrong order of slices. But, I'm afraid, faulty RAM won't get us anywhere > either because, as I have mentioned in the first message, I have two > backups like that, created 1 month apart, from two different systems > (different hardware, locations, age, usage patterns, etc. - only backup > mechanism and Ubuntu/dar versions are the same). Both backups report Even if it sounds improbable, the corruption by the RAM (of the system where dar runs) should be first invalidated. A test on one of the concerned nodes would be sufficient. So whenever you have the opportunity (production constraint, I can understand), a memtest86+ would worth running for a whole cycle (yes it may take a couple of hour...depending on how fast and how large is the RAM). In the past, dar already revealed RAM weaknesses that did not avoiding the OS to run normally... memtest86 brought the explanation. See also that some users make even bigger backups than what you seems doing, and so far (here I'm crossing my fingers!!!) no one reported any such data corruption... while 2.7.15 is already five years old. IMHO, there is something weird in your environment. I would also suggest in your backup process to leverage the --hash option of dar (--hash sha512 or --hash whirlpool at least), checking the hash right after each slice generation and before slice usage would let you completely eliminate storage and network as a cause of the problem, as the hash is performed in memory on the data just before it is written to disk. Though, the --hash option will not be able to reveal any problem regarding RAM issue.... I'm completely out of any other plausible track than memory/RAM issue, really... I really need you to test the RAM of one of your host where dar is running, before going further in investigations with more improbable hypothesis: eliminating the probable causes before considering the improbable ones... Cheers, Denis |
|
From: mannino <ma...@of...> - 2026-06-03 19:50:30
|
>> Other than a bug in dar itself (which I prefer to not consider thanks
>> to your exceptional job)
>
> exceptional or not, all software can have bugs and I never excluded it,
> the problem is the lack of reproducibility in your context...
That's true, until a new backup comes up with the same issue.
Full-system backups take a while to finish and they don't run every week
so let's stay tuned.
Meanwhile, you suggested we rule out disk errors, transfer errors and
wrong order of slices. But, I'm afraid, faulty RAM won't get us anywhere
either because, as I have mentioned in the first message, I have two
backups like that, created 1 month apart, from two different systems
(different hardware, locations, age, usage patterns, etc. - only backup
mechanism and Ubuntu/dar versions are the same). Both backups report
Sorry, file size is unknown at this step of the program.
to -l -0 -q; strace shows reading of the last slice, then some preceding
slice; file counts in -affs and -0 don't (fully) match:
dar -t -A -affs:
- 5689 x "compressed data CRC error"
- 1448 x "CRC error: data corruption"
- I haven't read through the entire list but it seems there are no
directories and corruptions affect files of various sizes from tiny to
huge and in different slices
--------------------------------------------
274471 item(s) treated
7137 item(s) with error
0 item(s) ignored (excluded by filters)
--------------------------------------------
Total number of items considered: 281608
--------------------------------------------
dar -t -0:
- 8 x "No escape mark found for that file"
--------------------------------------------
724494 item(s) treated
8 item(s) with error
0 item(s) ignored (excluded by filters)
--------------------------------------------
Total number of items considered: 724502
--------------------------------------------
dar -l <catalogue> -q:
Archive version format : 11.3
Compression algorithm used : bzip2
Compression block size used : 0
Symmetric key encryption used : none
Asymmetric key encryption used : none
Archive is signed : no
Sequential reading marks : present
User comment : redacted
Catalogue size in archive : 24538535 bytes
Archive is composed of 1 file(s)
File size: 24538747 bytes
The global data compression ratio is: 25%
WARNING! This archive only contains the catalogue of another archive,
it can only be used as reference for differential backup or as rescue in
case of corruption of the original archive's content. You cannot restore
any data from this archive alone
Archive of reference slicing:
Other slices: 2147483648 byte(s)
in-place path: redacted
CATALOGUE CONTENTS :
total number of inode : 724464
fully saved : 724464
binay delta patch : 0
inode metadata only : 0
distribution of inode(s)
- directories : 28009
- plain files : 687898
- symbolic links : 8403
- named pipes : 12
- unix sockets : 138
- character devices : 3
- block devices : 1
- Door entries : 0
hard links information
- number of inode with hard link : 21
- number of reference to hard linked inodes: 51
destroyed entries information
0 file(s) have been record as destroyed since backup of reference
This backup consists of 95 slices totaling 190 GiB.
Certainly it'd be too much of a coincidence if two servers exhibited
abnormal (RAM) behavior only during the creation of dar backups?
>> I still think it has something to do with wrong order/number of slices.
>
> to be honest, I don't see how mixing or even renumbering slices could
> cause the behavior you hit...
I only suggested that because dar doesn't detect out-of-sequence slices,
emitting symptomatic errors instead. If there were such a combination of
factors that could theoretically result in what I'm observing, I'd opt
for the mixed slices cause. But I surely trust you on this one... then
what else remains? What could potentially result in such archives?
>>> When testing the backup in sequential-read mode (-0/--sequential-read
>>> option) could you add the following option:
>>> -E "echo reading slice %N"
>>> and report the last slice being reported as being read by this added -E
>>> option, before the failure takes place?
>>
>> As per the above output, it says "reading slice 21".
>>
>>> Also could you tell me the number of the last slice available?
>>
>> There are arc.1.dar through arc.29.dar, the total of 29 slices.
>
> OK so you don't miss slices, but have those corrupted at different
> places.
But why does it report "Unexpected value while reading archive version"
after reading slice 21 after 29? It also happens with the other backup
except it reads slice 95 then 76.
> In sequential read mode, obviously, the archive is read in sequence,
> after having read the data of a files, dar reads right after the
> expected CRC and compares it with the one calculated on the retreived
> data and eventually reports a difference. Then it skips forward looking
> for the next tape mark and once again and the loop repeats...
>
> Here, corruption occurred in four different files (and slices), not only
> the very last files (in the last slice) as I had understood, which is
> not the pattern of a truncated slice.
It makes sense. It happens to the other backup too.
> Can you double check with for example memtest86+ ?
The system needs to be powered down for that, it's not something we can
do at this time...
>> 1. Archive UUID, as it's currently implemented, is perfect: (1) it
>> practically guarantees that one cannot mix up slices from different
>> archives; (2) dar itself checks that when testing; (3) can be
>> displayed with `file`:
>>
>> arc.10.dar: dar archive, label "82af71c2 00000000 fffff752"
>>
>> I could only suggest storing UUID of the reference archive in the -@
>> catalogue too (since -@ gets its own different UUID) and for dar
>> itself to provide some way to view this UUID. In any case, UUID helps
>> one work out the consistency among archives - but there's nothing
>> similar for working out slices for the archive itself, and it's a
>> problem for me.
>
> if you have a look at dar's documentation about the archive format, you
> will see that this is done with more flexibility:
> - internal name : links slices of an archive together
> - data name : links a catalog (or isolated catalog) to the data set to
> use. This prevents one trying to restore an archive with the help of an
> unrelated isolated catalog.
>
> http://dar.linux.free.fr/doc/Notes.html#archive_structure
I have glanced over that page but I have missed this. It's great and
eliminates any confusion in mismatched slices/catalogues. Now, if only
dar could provide some way to view this info, preferably in -0 mode for
potentially corrupted slices. For now, my only diagnostic tool is
`file`, and it's only reporting internal name it seems.
>> 2. Byte sizes of non/last slices would allow basic integrity check but
>> (1) `file` doesn't show them; (2) dar doesn't show them in -l -0 mode
>> (so unable to extract from invalid file); (3) dar -l -q works on
>> isolated catalogue but only reports non-last slice size - given our
>> problem here is likely last slice being truncated, it doesn't really
>> help.
>
> dar cannot know in advance how much slices will be necessary nor what
> will be the size of the last slice.
Of course it's streaming so I'm not suggesting storing total number of
slices in each slice. My suggestion here is to store each slice's size -
which is of course known beforehand. (I've written "non/last" but I
shall recede from the "last" part indeed.)
> The sar layer (which manages slices)
> processes a stream of bytes from upper layers (ciphering, compression,
> filtering, and so forth) which is generated on the fly from the
> filesystem to backup, the sar takes this flow and creates slices from it
> up to the time the flow dries up.
>
> What is stored at the beginning of *each slice* is the so called
> slice-layout:
> - first slice max size,
> - other slices max size,
> - slice header size
> - and slice format (2 formats so far).
>
> Only this could be shown in sequential-read mode, not what is displayed
> today in direct mode which adds to this the overall size of the archive
> and the size of the last slice (easy to get because in that mode this is
> the last slice that get read first).
Okay, the ability to view the details of slice-layout in -0 mode would
be welcome. I'd say any header info available to dar should be made
visible on equal terms.
>> 3. Slice number is what I truly miss. According to my tests, messing
>> with non-last slices (removing, duplicating, reordering) is usually
>> detected by dar as different types of data inconsistency so even if
>> you have all slices perfectly good and well, just in the wrong order,
>> you would think the backup is botched, and those errors would only be
>> misleading.
>
> the slice number is expected to be in the filename and it is not
> expected that users would change/renumber it...
Let me explain our storage strategy briefly. Backup files are held off
site for obvious reasons; the storage space is a set of buckets (could
be one or more per server but they're static) storing opaque blobs whose
only characteristics are acceptance date and data itself. Date is stored
for FIFO-like purging. Blobs can be moved or duplicated elsewhere behind
the scenes.
I understand it may not be the most elegant or efficient solution but
that's what we have now. Allowing free form metadata like file names
would complicate blob management. For backup purposes, this is not a
problem because each backup location is tied to a particular bucket
where archives and slices go in chronological order so by inspecting
date and data size one can clearly divide entries into individual
archives even if there are no file names. We store additional data in
dar's user comment field, but it's obviously archive-wise.
So essentially our backup directory looks like this:
402 | Sun Feb 22 11:43:58 AM UTC 2026 | 167511621 bytes
403 | Tue Mar 3 06:56:46 PM UTC 2026 | 10613869728 bytes
404 | Tue Mar 3 07:13:29 PM UTC 2026 | 10613869728 bytes
405 | Tue Mar 3 07:36:56 PM UTC 2026 | 635571924 bytes
> This is something unclear to me: you have this slice number information
> in the slice filename... isn't it the easiest way to know what slice
> number is a slice? And how much slice do one have? And also what is the
> latest slice?
My point is that in some scenarios file names are unavailable. Even if
not for our storage system, file names may be missing as a result of
data recovery from a faulty drive, for example. Sure, that should not
normally happen, and if you use par2 then it would fix the names for you
as well, but we're talking about last-chance recovery here, and that's
why dar has the -0 mode, so why not take it a step further with one or
two simple new fields in the header?
> What I don't understand is how you could
> accidentally miss the first slice, I mean how do you let dar accessing
> the slices as dar, even in sequential-read mode open a slice after a
> given filename and if it is not present it complains and waits.
To deploy or even inspect a backup, its blobs should be pulled from
storage somewhere dar can run, then renamed or symlinked to form
arc.N.dar. Human factor makes such mistakes possible. If you forgot one
or more leading slices yet they're named arc.1.dar, arc.2.dar, etc. -
dar won't "complain and wait", it'd proceed straight to throwing errors
at you.
>> Also, out of curiosity, I tried copying one slice twice and got a ton
>> of CRC errors.
> > > With the help of 4 extra bytes (= slice number) in each slice's
> header,
>> we would be able to identify the problem precisely.
>
> why only 4 bytes? This would limit the number of slices... even if today
> you think it is large enough, this is the same way of thinking that led
> to the year 2000 bug, the problem of year 2038, the upper and high
> memory in DOS some decades ago, the 2 GB files boundary (and need of
> large file support)... and so forth.
I understand your point. But even if we take 1 KiB per slice, 2^32
slices would add up to 4.4 TiB of data... and more "realistic" 1 MiB per
slice would be 4.5 PiB... but if that is a concern, you could use 64
bits, that would be (let me look that up on Wikipedia) around 2 exabytes
of data if each slice only held 1 byte of payload... I'm afraid we'd
need more than quantum computers to cross that boundary...
> A better approach could be to add a random fixed size field at end of
> slice and have this same random value at the beginning of the next. But
> I don't see the use case as mentioned earlier about number of slices in
> filenames.
If we're in the territory where 2^32 slices may not be enough, our
random numbers may not be random enough too, especially within millions
of slices... If we were choosing, I'd just use a fixed length numeric
field for the sake of simplicity. VLQ if you must, although I'd grade
that as an overkill...
>> Not only will it tell us slice order, it will also tell us the total
>> number of slices (since identifying last slice is trivial).
>
> not to my standpoint, as all slices may not be visible at a given time
> in particular when using removal media or remote storage.
Again, I'm not advocating storing *total* number of slices anywhere. My
point is that, with a simple "current slice index" field embedded in
each slice-layout, it's possible to eliminate all slice order-related
errors and usually even derive the total number of slices by taking last
slice's counter (given that last slice is smaller than others, detecting
it empirically is trivial). Sparing several bytes in each slice is a
small price to pay for this benefit, IMHO.
>> Storing total number of slices in the catalogue would take it even
>> further,
>
> this would break the ability to re-slice (dar_xform) an archive, once
> the isolated catalogue would have been created.
Okay, no worries, I take this suggestion back.
Thank you for still being on board with me on this!
|
|
From: Denis C. <dar...@fr...> - 2026-06-03 13:59:15
|
Le 02/06/2026 à 23:05, mannino a écrit : > As usual, thank you for responding, Denis! you are welcome [...] > Data corruption is unlikely to happen on two different systems to just > the header block out of all other 100s of GiBs of dar data. I fully agree, this is weird... > Also, since > the corruption is in the last slice (?), last slice is never larger than > other slices, each slice is transferred to another system and deleted (- > E) and non-last slices were not truncated, free space shortage shouldn't > be the cause either. correct > > Other than a bug in dar itself (which I prefer to not consider thanks to > your exceptional job) exceptional or not, all software can have bugs and I never excluded it, the problem is the lack of reproducibility in your context... But as you mentioned you are not using a complex mix of features from dar and those you are using are broadly used and tested: there are non-regression tests running several hundred thousands times dar, mixing a lot of feature on and off and taking weeks to complete, that is run fully at least for each major release... So there may still be a bug that express from time to time, but I suspect it is somehow also strongly related to your context... >, I still think it has something to do with wrong > order/number of slices. to be honest, I don't see how mixing or even renumbering slices could cause the behavior you hit... > But I cannot be certain - validating this (or > any other) theory is hard without enough metadata: yep. > > 1. Archive UUID, as it's currently implemented, is perfect: (1) it > practically guarantees that one cannot mix up slices from different > archives; (2) dar itself checks that when testing; (3) can be displayed > with `file`: > > arc.10.dar: dar archive, label "82af71c2 00000000 fffff752" > > I could only suggest storing UUID of the reference archive in the -@ > catalogue too (since -@ gets its own different UUID) and for dar itself > to provide some way to view this UUID. In any case, UUID helps one work > out the consistency among archives - but there's nothing similar for > working out slices for the archive itself, and it's a problem for me. if you have a look at dar's documentation about the archive format, you will see that this is done with more flexibility: - internal name : links slices of an archive together - data name : links a catalog (or isolated catalog) to the data set to use. This prevents one trying to restore an archive with the help of an unrelated isolated catalog. http://dar.linux.free.fr/doc/Notes.html#archive_structure But yes, it could be shown when listing archive details... > > 2. Byte sizes of non/last slices would allow basic integrity check but > (1) `file` doesn't show them; (2) dar doesn't show them in -l -0 mode > (so unable to extract from invalid file); (3) dar -l -q works on > isolated catalogue but only reports non-last slice size - given our > problem here is likely last slice being truncated, it doesn't really help. dar cannot know in advance how much slices will be necessary nor what will be the size of the last slice. The sar layer (which manages slices) processes a stream of bytes from upper layers (ciphering, compression, filtering, and so forth) which is generated on the fly from the filesystem to backup, the sar takes this flow and creates slices from it up to the time the flow dries up. What is stored at the beginning of *each slice* is the so called slice-layout: - first slice max size, - other slices max size, - slice header size - and slice format (2 formats so far). Only this could be shown in sequential-read mode, not what is displayed today in direct mode which adds to this the overall size of the archive and the size of the last slice (easy to get because in that mode this is the last slice that get read first). > > 3. Slice number is what I truly miss. According to my tests, messing > with non-last slices (removing, duplicating, reordering) is usually > detected by dar as different types of data inconsistency so even if you > have all slices perfectly good and well, just in the wrong order, you > would think the backup is botched, and those errors would only be > misleading. the slice number is expected to be in the filename and it is not expected that users would change/renumber it... > > For example, on one of my backups (40+ slices, 200 GiB) I accidentally > missed the first slice. dar -t reported: > > Reached End of File while reading archive version > > dar -t -0: > > Unexpected value while reading archive version yes, both way open the archive in a different way, so the inconsistency is reported differently. What I don't understand is how you could accidentally miss the first slice, I mean how do you let dar accessing the slices as dar, even in sequential-read mode open a slice after a given filename and if it is not present it complains and waits. > > Also, out of curiosity, I tried copying one slice twice and got a ton of > CRC errors. > > With the help of 4 extra bytes (= slice number) in each slice's header, > we would be able to identify the problem precisely. why only 4 bytes? This would limit the number of slices... even if today you think it is large enough, this is the same way of thinking that led to the year 2000 bug, the problem of year 2038, the upper and high memory in DOS some decades ago, the 2 GB files boundary (and need of large file support)... and so forth. A better approach could be to add a random fixed size field at end of slice and have this same random value at the beginning of the next. But I don't see the use case as mentioned earlier about number of slices in filenames. > Not only will it > tell us slice order, it will also tell us the total number of slices > (since identifying last slice is trivial). not to my standpoint, as all slices may not be visible at a given time in particular when using removal media or remote storage. > Storing total number of > slices in the catalogue would take it even further, this would break the ability to re-slice (dar_xform) an archive, once the isolated catalogue would have been created. > but even just > counters would be an incredible step forward for me. This is something unclear to me: you have this slice number information in the slice filename... isn't it the easiest way to know what slice number is a slice? And how much slice do one have? And also what is the latest slice? By the way, you even have the --min-digit option to have slices sorted in lexical order... my 2 cents... > [...] >> >> The good thing however is that you should be able to restore 1619434 >> files and directories from your backup using the --sequential-read mode. > > Yes, the recovery went fine. OK, good to know! > Comparing it with the isolated catalogue > that was created on the fly when the backup was made, it seems > everything apart from those 4 files was restored correctly. My main > issue here isn't recovery of these particular archives per se, it's > trying to find how exactly they are corrupted so I can go on finding the > cause and avoid this in the future. sure! [...] >> >> When testing the backup in sequential-read mode (-0/--sequential-read >> option) could you add the following option: >> -E "echo reading slice %N" >> and report the last slice being reported as being read by this added -E >> option, before the failure takes place? > > As per the above output, it says "reading slice 21". > >> Also could you tell me the number of the last slice available? > > There are arc.1.dar through arc.29.dar, the total of 29 slices. OK so you don't miss slices, but have those corrupted at different places. > >> > > Does it indicate that the archive is fine sans for 4 files' data? >> Then >>> what's with that "Unexpected value" error in direct mode? >> >> it indicates that before the data corruption (= Unexpected value) you >> can restore 1619434 files and directories and 4 entries (probably 4 >> nested directories opened but not closed (?)) are reported as having an >> error. > > No, those entries correspond to regular files (some tiny, some huge) > spread among different slices. OK, that's interesting! > >> Doesn't dar report the entry names of that 4 failed entries??? > > It does, I just redacted it - sorry for the confusion. They're all in > different slices and directories so it shouldn't really matter. In sequential read mode, obviously, the archive is read in sequence, after having read the data of a files, dar reads right after the expected CRC and compares it with the one calculated on the retreived data and eventually reports a difference. Then it skips forward looking for the next tape mark and once again and the loop repeats... Here, corruption occurred in four different files (and slices), not only the very last files (in the last slice) as I had understood, which is not the pattern of a truncated slice. [...] > >> Do you still have the isolated catalog of the full backups? >> >> If so you could try the following command: >> dar -t <backup> -A <isolated cat> -affs >> >> the -affs is necessary because the last slice is obviously corrupted, >> this lead dar to read the archive header format and other important >> stuff from the copy located in the first slice. This option/feature was >> added in 2.8.x. >> >> Doing so will give you an exact picture of the amount of data you are >> missing from your backup. > > dar -t -A -affs: > > - 790 x CRC error: data corruption > - 3046 x compressed data CRC error > > -------------------------------------------- > 715129 item(s) treated > 3836 item(s) with error > 0 item(s) ignored (excluded by filters) > -------------------------------------------- > Total number of items considered: 718965 > -------------------------------------------- > 718965 inodes tested, 790 with CRC error (not compressed ones) and 3046 with error in data compressed structure (compressed ones)... I would really double check the RAM... but read further: > > dar -t -0: > > - 4 x can't read data CRC: No escape mark found for that file > > -------------------------------------------- > 1619434 item(s) treated > 4 item(s) with error > 0 item(s) ignored (excluded by filters) > -------------------------------------------- > Total number of items considered: 1649928 > -------------------------------------------- > It is possible to explain the different number of faults on CRC between the two reading modes, due to the fact this is not the same CRC copy that is used to compare with the data (interleaved with data in sequential-read mode, stored in memory during the backup process and wrote to disk at the end in the form of a catalogue, in direct mode). But I cannot understand why you get 3046 compressed data error on one side only, unless the compression algorithm used has been tampered in memory (not on storage because then it would have lead to a CRC mismatch when loading the catalogue to memory). Now about the total number of items... Accordingly to the details report you provided from the isolated catalogue, I can explain almost the 1649928 entries: it corresponds to: total number of inode : 1589967 plus number of reference to hard linked inodes: 109555 but less: number of inode with hard link : 49600 because each first time an hard linked inode is met is also the first time a hard link to this item is met and both counts as 1, then subsequent hard links count also 1. So you must avoid counting twice the first hard link, but need to add to the inodes count those of hard-links found in the catalogue content summary. This makes 1649922, though I can't explain the difference of 6 files, but I double checked the code, did some tests and get this expected counts and got the same in both sequential and direct modes and even when manually corrupting the archive file. Now I cannot explain why the same isolated catalogue when applied to the corrupted archive as backup only report 718965 entries, this is really weird. They should all be tried in turn and lead to increase either the pass or fail counter, for a total count to match the one found in sequential read mode. Thus, I would really double check the RAM, not the storage, because if there was a corruption at storage level, the catalogue would not have been loaded properly and an error would have shown. Can you double check with for example memtest86+ ? > > dar -t <catalogue>: > > WARNING! This is an isolated catalogue, no data or EA is present in > this archive, only the catalogue structure can be checked > > -------------------------------------------- > 0 item(s) treated > 0 item(s) with error > 0 item(s) ignored (excluded by filters) > -------------------------------------------- > Total number of items considered: 0 > -------------------------------------------- > > > dar -l <catalogue> -q: > > Archive version format : 11.3 > Compression algorithm used : bzip2 > Compression block size used : 0 > Symmetric key encryption used : none > Asymmetric key encryption used : none > Archive is signed : no > Sequential reading marks : present > User comment : > Catalogue size in archive : 35328204 bytes > > Archive is composed of 1 file(s) > File size: 35328416 bytes > The global data compression ratio is: 37% > > WARNING! This archive only contains the catalogue of another archive, > it can only be used as reference for differential backup or as rescue in > case of corruption of the original archive's content. You cannot restore > any data from this archive alone > > Archive of reference slicing: > Other slices: 10737418240 byte(s) > > in-place path: redacted > > CATALOGUE CONTENTS : > > total number of inode : 1589967 > fully saved : 1589967 > binay delta patch : 0 > inode metadata only : 0 > distribution of inode(s) > - directories : 236542 > - plain files : 1322581 > - symbolic links : 30513 > - named pipes : 16 > - unix sockets : 33 > - character devices : 280 > - block devices : 2 > - Door entries : 0 > hard links information > - number of inode with hard link : 49600 > - number of reference to hard linked inodes: 109555 > destroyed entries information > 0 file(s) have been record as destroyed since backup of reference > > > file arc.cat.1.dar arc.1.dar: > > arc.cat.1.dar: dar archive, label "7929f969 00000000 8a33" end slice > arc.1.dar: dar archive, label "0503f969 00000000 8a33" > > > Why -affs, -0, -l -q <cat> all report different file counts? good question, as detailed above. > |
|
From: mannino <ma...@of...> - 2026-06-02 21:06:50
|
As usual, thank you for responding, Denis!
I have conducted some tests but, as it often happens, an attempt to
explicitly repro an issue goes without a hassle. Guess I will have to
keep an eye on new backups to see if this resurfaces. Meanwhile, I still
have some questions/suggestions for retrospective analysis.
>> I have two different systems that back up data using incremental
>> archives. Those incremental archives (there's about 200 of them) are
>> fine according to dar -t. However, base (full) archives produced on
>> both systems at different times are reported as corrupted.
>>
>> Adding -va to -t provides little insight into what's going on:
>>
>> # dar version 2.7.17
>>
>> Auto detecting min-digits to be 2
>> Opening archive aa ...
>> Opening the archive using the multi-slice abstraction layer...
>> Reading the archive trailer...
>> Final memory cleanup...
>> FATAL error, aborting operation: Unexpected value while reading
>> archive version
>
> this means that the format version field is not properly formatted and
> dar cannot know which is the internal structure of the archive.
Data corruption is unlikely to happen on two different systems to just
the header block out of all other 100s of GiBs of dar data. Also, since
the corruption is in the last slice (?), last slice is never larger than
other slices, each slice is transferred to another system and deleted
(-E) and non-last slices were not truncated, free space shortage
shouldn't be the cause either.
Other than a bug in dar itself (which I prefer to not consider thanks to
your exceptional job), I still think it has something to do with wrong
order/number of slices. But I cannot be certain - validating this (or
any other) theory is hard without enough metadata:
1. Archive UUID, as it's currently implemented, is perfect: (1) it
practically guarantees that one cannot mix up slices from different
archives; (2) dar itself checks that when testing; (3) can be displayed
with `file`:
arc.10.dar: dar archive, label "82af71c2 00000000 fffff752"
I could only suggest storing UUID of the reference archive in the -@
catalogue too (since -@ gets its own different UUID) and for dar itself
to provide some way to view this UUID. In any case, UUID helps one work
out the consistency among archives - but there's nothing similar for
working out slices for the archive itself, and it's a problem for me.
2. Byte sizes of non/last slices would allow basic integrity check but
(1) `file` doesn't show them; (2) dar doesn't show them in -l -0 mode
(so unable to extract from invalid file); (3) dar -l -q works on
isolated catalogue but only reports non-last slice size - given our
problem here is likely last slice being truncated, it doesn't really help.
3. Slice number is what I truly miss. According to my tests, messing
with non-last slices (removing, duplicating, reordering) is usually
detected by dar as different types of data inconsistency so even if you
have all slices perfectly good and well, just in the wrong order, you
would think the backup is botched, and those errors would only be
misleading.
For example, on one of my backups (40+ slices, 200 GiB) I accidentally
missed the first slice. dar -t reported:
Reached End of File while reading archive version
dar -t -0:
Unexpected value while reading archive version
Also, out of curiosity, I tried copying one slice twice and got a ton of
CRC errors.
With the help of 4 extra bytes (= slice number) in each slice's header,
we would be able to identify the problem precisely. Not only will it
tell us slice order, it will also tell us the total number of slices
(since identifying last slice is trivial). Storing total number of
slices in the catalogue would take it even further, but even just
counters would be an incredible step forward for me.
>> I tried checking with strace to see what files dar -t opens. For both
>> archives, it starts with the last one in the directory, then opens
>> some seemingly random one in the 2/3rd of the set (I have ~30 and ~100
>> slices so it opens the ~30th/~70th respectively), then after a bunch
>> of seeks it fails.
>
> Weird, the archive version information is fetched very early in the
> archive reading process, if this field is corrupted, dar should open the
> last slice only and fail with the message you reported.
Yet this is exactly what happens:
# strace -efile dar -t arc -E "echo reading slice %N"
...
openat(AT_FDCWD, ".", O_RDONLY|O_NONBLOCK|O_CLOEXEC|O_DIRECTORY) = 4
reading slice 29
--- SIGCHLD {si_signo=SIGCHLD, si_code=CLD_EXITED, si_pid=19001,
si_uid=0, si_status=0, si_utime=0, si_stime=0} ---
openat(AT_FDCWD, "/tmp/arc.29.dar", O_RDONLY) = 4
reading slice 21
--- SIGCHLD {si_signo=SIGCHLD, si_code=CLD_EXITED, si_pid=19002,
si_uid=0, si_status=0, si_utime=0, si_stime=0} ---
openat(AT_FDCWD, "/tmp/arc.21.dar", O_RDONLY) = 4
openat(AT_FDCWD, "/usr/share/locale/en_US.UTF-8/LC_MESSAGES/dar.mo",
O_RDONLY) = -1 ENOENT (No such file or directory)
<locale stuff redacted>
Final memory cleanup...
FATAL error, aborting operation: Unexpected value while reading
archive version
+++ exited with 2 +++
>> I tried -t -0. It resulted in an entry like so for 4 files:
>>
>> can't read data CRC: No escape mark found for that file
>>
>> ...followed by this output:
>>
>> --------------------------------------------
>> 1619434 item(s) treated
>> 4 item(s) with error
>> 0 item(s) ignored (excluded by filters)
>> --------------------------------------------
>> Total number of items considered: 1619438
>> --------------------------------------------
>> Final memory cleanup...
>> Some files are corrupted in the archive and it will not be possible
>> to restore them
>> +++ exited with 5 +++
>
> the final error is probably because the archive has been interrupted or
> because the disk was full when it was created and dar was not in
> interactive mode (it was run from as a cron job, for example).
>
> This would also explain why without -0 (--sequential-read) option dar is
> not able to find the archive version within a proper format, because
> this information is lacking at the end of the last slice.
>
> But the data at the end of the archive seems having a normal terminator
> and slice trailer is present (else the error would have been different)...
>
> The good thing however is that you should be able to restore 1619434
> files and directories from your backup using the --sequential-read mode.
Yes, the recovery went fine. Comparing it with the isolated catalogue
that was created on the fly when the backup was made, it seems
everything apart from those 4 files was restored correctly. My main
issue here isn't recovery of these particular archives per se, it's
trying to find how exactly they are corrupted so I can go on finding the
cause and avoid this in the future.
>> No other errors were logged. None were logged for the first and last
>> slices.
>
> When testing the backup in sequential-read mode (-0/--sequential-read
> option) could you add the following option:
> -E "echo reading slice %N"
> and report the last slice being reported as being read by this added -E
> option, before the failure takes place?
As per the above output, it says "reading slice 21".
> Also could you tell me the number of the last slice available?
There are arc.1.dar through arc.29.dar, the total of 29 slices.
> > > Does it indicate that the archive is fine sans for 4 files' data? Then
>> what's with that "Unexpected value" error in direct mode?
>
> it indicates that before the data corruption (= Unexpected value) you
> can restore 1619434 files and directories and 4 entries (probably 4
> nested directories opened but not closed (?)) are reported as having an
> error.
No, those entries correspond to regular files (some tiny, some huge)
spread among different slices.
> Doesn't dar report the entry names of that 4 failed entries???
It does, I just redacted it - sorry for the confusion. They're all in
different slices and directories so it shouldn't really matter.
>> -l -q -0: fails with (???):
>>
>> Sorry, file size is unknown at this step of the program.
>> The last file of the set is not present in file:///, please provide
>> it.
>
> interesting, this gives weight to the scenario of the truncated backup...
>
>>
>> Finally, I tried -al to no avail:
>>
>> LAX MODE: Failed to read the archive format version.
>> LAX MODE: Please provide the format MAJOR number: 11
>> LAX MODE: Please provide the format MINOR number: 3
>> LAX MODE: Using archive format "11.3"? [return = YES | Esc = NO]
>> Continuing...
>
> OK so we dar continues as if the format field was correct
>
>> LAX MODE: Unknown compression algorithm used, assuming data
>> corruption occurred. Please help me, answering with one of the
>> following words "none", "gzip", [...] at the next prompt:gzip
>
> This means that not only the archive format field is corrupted but also
> the compression algo field...
>
>> Unknown crypto algorithm used in archive, ignoring that field and
>> simply assuming the archive has been encrypted, if not done you will
>> need to specify the crypto algorithm to use in order to read this archive
>> Error met while reading archive of reference slicing layout,
>> ignoring this field and continuing
>> Aborting due to exception: Badly formed "infinint" or not supported
>> format
>
> Here you could get one step further adding '-K none:' on command-line to
> tell dar that the archive was not encrypted (if it was not, of course).
It wasn't encrypted. Adding -K none: didn't change anything except the
"Unknown crypto algorithm" line disappeared:
...
LAX MODE: Unknown compression algorithm used, assuming data
corruption occurred. Please help me, answering with one of the following
words "none", "gzip", "bzip2", "lzo", "xz", "zstd" or "lz4" at the next
prompt:gzip
Error met while reading archive of reference slicing layout, ignoring
this field and continuing
Final memory cleanup...
FATAL error, aborting operation: Badly formed "infinint" or not
supported format
> Do you still have the isolated catalog of the full backups?
>
> If so you could try the following command:
> dar -t <backup> -A <isolated cat> -affs
>
> the -affs is necessary because the last slice is obviously corrupted,
> this lead dar to read the archive header format and other important
> stuff from the copy located in the first slice. This option/feature was
> added in 2.8.x.
>
> Doing so will give you an exact picture of the amount of data you are
> missing from your backup.
dar -t -A -affs:
- 790 x CRC error: data corruption
- 3046 x compressed data CRC error
--------------------------------------------
715129 item(s) treated
3836 item(s) with error
0 item(s) ignored (excluded by filters)
--------------------------------------------
Total number of items considered: 718965
--------------------------------------------
dar -t -0:
- 4 x can't read data CRC: No escape mark found for that file
--------------------------------------------
1619434 item(s) treated
4 item(s) with error
0 item(s) ignored (excluded by filters)
--------------------------------------------
Total number of items considered: 1649928
--------------------------------------------
dar -t <catalogue>:
WARNING! This is an isolated catalogue, no data or EA is present in
this archive, only the catalogue structure can be checked
--------------------------------------------
0 item(s) treated
0 item(s) with error
0 item(s) ignored (excluded by filters)
--------------------------------------------
Total number of items considered: 0
--------------------------------------------
dar -l <catalogue> -q:
Archive version format : 11.3
Compression algorithm used : bzip2
Compression block size used : 0
Symmetric key encryption used : none
Asymmetric key encryption used : none
Archive is signed : no
Sequential reading marks : present
User comment :
Catalogue size in archive : 35328204 bytes
Archive is composed of 1 file(s)
File size: 35328416 bytes
The global data compression ratio is: 37%
WARNING! This archive only contains the catalogue of another archive,
it can only be used as reference for differential backup or as rescue in
case of corruption of the original archive's content. You cannot restore
any data from this archive alone
Archive of reference slicing:
Other slices: 10737418240 byte(s)
in-place path: redacted
CATALOGUE CONTENTS :
total number of inode : 1589967
fully saved : 1589967
binay delta patch : 0
inode metadata only : 0
distribution of inode(s)
- directories : 236542
- plain files : 1322581
- symbolic links : 30513
- named pipes : 16
- unix sockets : 33
- character devices : 280
- block devices : 2
- Door entries : 0
hard links information
- number of inode with hard link : 49600
- number of reference to hard linked inodes: 109555
destroyed entries information
0 file(s) have been record as destroyed since backup of reference
file arc.cat.1.dar arc.1.dar:
arc.cat.1.dar: dar archive, label "7929f969 00000000 8a33" end slice
arc.1.dar: dar archive, label "0503f969 00000000 8a33"
Why -affs, -0, -l -q <cat> all report different file counts?
|
|
From: Denis C. <dar...@fr...> - 2026-05-27 12:23:56
|
Le 27/05/2026 à 00:47, mannino a écrit : > I have two different systems that back up data using incremental > archives. Those incremental archives (there's about 200 of them) are > fine according to dar -t. However, base (full) archives produced on both > systems at different times are reported as corrupted. > > Adding -va to -t provides little insight into what's going on: > > # dar version 2.7.17 > > Auto detecting min-digits to be 2 > Opening archive aa ... > Opening the archive using the multi-slice abstraction layer... > Reading the archive trailer... > Final memory cleanup... > FATAL error, aborting operation: Unexpected value while reading > archive version this means that the format version field is not properly formatted and dar cannot know which is the internal structure of the archive. > > # dar version 2.7.15 > > Opening archive aa ... > Opening the archive using the multi-slice abstraction layer... > Reading the archive trailer... > Final memory cleanup... > FATAL error, aborting operation: Unexpected value while reading > archive version same thing here > > When creating archives, I don't use any fancy options, just -zgzip, -A/- > @ and -s/-E. I do have isolated catalogues in addition to the archives > themselves but adding -A to -t or -l doesn't change anything. I have double checked the Changelog since version 2.7.17 to 2.7.21 (latest of branch 2.7) and could not find any known bug that could explain this behavior, so its OK to keep 2.7.17 for further investigations. And you are not using complicated mix of options that could have gone through non-regression tests without a problem being reported. > > Archives are stored on a separate machine; it's possible that a > corruption indeed took place when those archives were being transferred > from hosts but those hosts are different enough to make it unlikely that > that would happen to both at once. that's a possibility: you can eliminate this possible cause by transferring --- exactly as you did for the backups --- any *binary* file (like /bin/bash for example) and transferring it back to the hosts *beside* its original copy (but do not overwrite it!!!), then check with 'diff -s <original> <copy>' that the file are reported as identical. > > I tried checking with strace to see what files dar -t opens. For both > archives, it starts with the last one in the directory, then opens some > seemingly random one in the 2/3rd of the set (I have ~30 and ~100 slices > so it opens the ~30th/~70th respectively), then after a bunch of seeks > it fails. Weird, the archive version information is fetched very early in the archive reading process, if this field is corrupted, dar should open the last slice only and fail with the message you reported. > > I tried -t -0. It resulted in an entry like so for 4 files: > > can't read data CRC: No escape mark found for that file > > ...followed by this output: > > -------------------------------------------- > 1619434 item(s) treated > 4 item(s) with error > 0 item(s) ignored (excluded by filters) > -------------------------------------------- > Total number of items considered: 1619438 > -------------------------------------------- > Final memory cleanup... > Some files are corrupted in the archive and it will not be possible > to restore them > +++ exited with 5 +++ the final error is probably because the archive has been interrupted or because the disk was full when it was created and dar was not in interactive mode (it was run from as a cron job, for example). This would also explain why without -0 (--sequential-read) option dar is not able to find the archive version within a proper format, because this information is lacking at the end of the last slice. But the data at the end of the archive seems having a normal terminator and slice trailer is present (else the error would have been different)... The good thing however is that you should be able to restore 1619434 files and directories from your backup using the --sequential-read mode. > > No other errors were logged. None were logged for the first and last > slices. When testing the backup in sequential-read mode (-0/--sequential-read option) could you add the following option: -E "echo reading slice %N" and report the last slice being reported as being read by this added -E option, before the failure takes place? Also could you tell me the number of the last slice available? > > Does it indicate that the archive is fine sans for 4 files' data? Then > what's with that "Unexpected value" error in direct mode? it indicates that before the data corruption (= Unexpected value) you can restore 1619434 files and directories and 4 entries (probably 4 nested directories opened but not closed (?)) are reported as having an error. Doesn't dar report the entry names of that 4 failed entries??? But pay attention: if your backup has been truncated or disk space was lacking (and you did not pay attention to the error message/exit code dar reported), you may lack much more files that could not be saved. > > I also tried some combinations of -l and it really puzzled me. > > -l alone; -l -q: fail just like -t. this is normal, seen the reported error (corrupted archive format field) > > -l -0: after some output, fails with Error while listing archive > contents: can't read data CRC: No escape mark found for that file normal, it follows the same path as '-t -0' and fails for the same reason. The only difference is that it does not check crc of data and EA but skips over to reach the next metadata interleaved with the data, using so called "escape marks". > > -l -q -0: fails with (???): > > Sorry, file size is unknown at this step of the program. > The last file of the set is not present in file:///, please provide it. interesting, this gives weight to the scenario of the truncated backup... > > Finally, I tried -al to no avail: > > LAX MODE: Failed to read the archive format version. > LAX MODE: Please provide the format MAJOR number: 11 > LAX MODE: Please provide the format MINOR number: 3 > LAX MODE: Using archive format "11.3"? [return = YES | Esc = NO] > Continuing... OK so we dar continues as if the format field was correct > LAX MODE: Unknown compression algorithm used, assuming data > corruption occurred. Please help me, answering with one of the following > words "none", "gzip", [...] at the next prompt:gzip This means that not only the archive format field is corrupted but also the compression algo field... > Unknown crypto algorithm used in archive, ignoring that field and > simply assuming the archive has been encrypted, if not done you will > need to specify the crypto algorithm to use in order to read this archive > Error met while reading archive of reference slicing layout, ignoring > this field and continuing > Aborting due to exception: Badly formed "infinint" or not supported > format Here you could get one step further adding '-K none:' on command-line to tell dar that the archive was not encrypted (if it was not, of course). > > One plausible cause is that I have mixed up the slices (lost or > duplicated some intermediate slices). It seems dar header can't help me > here (it doesn't store slice index) but other evidence suggests it's not > the case (`file *.dar` outputs the same label, slices' file sizes and > times look sane, etc.). When you have mixed slices of different backups under the same basename, dar reports that. In each slice header is present an 'internal name' which is a random string based on the current time (UTC) computed at backup creation time. There is little chance (not to say no chance at all), that two different backups have the same internal name. From the first slice opened, dar record this internal name then compares it with subsequent slices it opens for that backup, if it does not match, dar stop and reports this inconsistency. > > Is -t -0 guaranteed to catch the situation when slices are mixed up? yes, this works the same (as explained just above) with and without -0 option. And also, files will not be restored if the embedded CRC in the archive does not match the expected one (also stored in the archive). > > For good measure, I tried -t'esting with dar 2.7.x and 2.8.x but results > were the same. > > I really don't know what to make of this or what else to try. You should be confident on the data you can restore from the backup when you use -0 option, but you must consider some data (maybe more than 'some' and probably more than the 4 failed reported entries) are missing from this "full" backup. Do you still have the isolated catalog of the full backups? If so you could try the following command: dar -t <backup> -A <isolated cat> -affs the -affs is necessary because the last slice is obviously corrupted, this lead dar to read the archive header format and other important stuff from the copy located in the first slice. This option/feature was added in 2.8.x. Doing so will give you an exact picture of the amount of data you are missing from your backup. > > Any help on finding the cause for corruption is much appreciated. > As you did differential backups after the full backup, it means that at that time the full backup was sane, else you would not have been able to create an isolated catalogue or to directly use the full backup as reference to build such differential/incremental backup, well except if you did an on-fly isolation and could find some available storage to have it written down. If you are lucky, the missing files are not so much and/or most of them have a more recent version available somewhere in a differential backup. So you wish can get back most of your data. Regards, Denis |
|
From: mannino <ma...@of...> - 2026-05-26 23:06:20
|
I have two different systems that back up data using incremental archives. Those incremental archives (there's about 200 of them) are fine according to dar -t. However, base (full) archives produced on both systems at different times are reported as corrupted. Adding -va to -t provides little insight into what's going on: # dar version 2.7.17 Auto detecting min-digits to be 2 Opening archive aa ... Opening the archive using the multi-slice abstraction layer... Reading the archive trailer... Final memory cleanup... FATAL error, aborting operation: Unexpected value while reading archive version # dar version 2.7.15 Opening archive aa ... Opening the archive using the multi-slice abstraction layer... Reading the archive trailer... Final memory cleanup... FATAL error, aborting operation: Unexpected value while reading archive version When creating archives, I don't use any fancy options, just -zgzip, -A/-@ and -s/-E. I do have isolated catalogues in addition to the archives themselves but adding -A to -t or -l doesn't change anything. Archives are stored on a separate machine; it's possible that a corruption indeed took place when those archives were being transferred from hosts but those hosts are different enough to make it unlikely that that would happen to both at once. I tried checking with strace to see what files dar -t opens. For both archives, it starts with the last one in the directory, then opens some seemingly random one in the 2/3rd of the set (I have ~30 and ~100 slices so it opens the ~30th/~70th respectively), then after a bunch of seeks it fails. I tried -t -0. It resulted in an entry like so for 4 files: can't read data CRC: No escape mark found for that file ...followed by this output: -------------------------------------------- 1619434 item(s) treated 4 item(s) with error 0 item(s) ignored (excluded by filters) -------------------------------------------- Total number of items considered: 1619438 -------------------------------------------- Final memory cleanup... Some files are corrupted in the archive and it will not be possible to restore them +++ exited with 5 +++ No other errors were logged. None were logged for the first and last slices. Does it indicate that the archive is fine sans for 4 files' data? Then what's with that "Unexpected value" error in direct mode? I also tried some combinations of -l and it really puzzled me. -l alone; -l -q: fail just like -t. -l -0: after some output, fails with Error while listing archive contents: can't read data CRC: No escape mark found for that file -l -q -0: fails with (???): Sorry, file size is unknown at this step of the program. The last file of the set is not present in file:///, please provide it. Finally, I tried -al to no avail: LAX MODE: Failed to read the archive format version. LAX MODE: Please provide the format MAJOR number: 11 LAX MODE: Please provide the format MINOR number: 3 LAX MODE: Using archive format "11.3"? [return = YES | Esc = NO] Continuing... LAX MODE: Unknown compression algorithm used, assuming data corruption occurred. Please help me, answering with one of the following words "none", "gzip", [...] at the next prompt:gzip Unknown crypto algorithm used in archive, ignoring that field and simply assuming the archive has been encrypted, if not done you will need to specify the crypto algorithm to use in order to read this archive Error met while reading archive of reference slicing layout, ignoring this field and continuing Aborting due to exception: Badly formed "infinint" or not supported format One plausible cause is that I have mixed up the slices (lost or duplicated some intermediate slices). It seems dar header can't help me here (it doesn't store slice index) but other evidence suggests it's not the case (`file *.dar` outputs the same label, slices' file sizes and times look sane, etc.). Is -t -0 guaranteed to catch the situation when slices are mixed up? For good measure, I tried -t'esting with dar 2.7.x and 2.8.x but results were the same. I really don't know what to make of this or what else to try. Any help on finding the cause for corruption is much appreciated. |
|
From: Denis C. <dar...@fr...> - 2026-04-21 19:22:09
|
Le 21/04/2026 à 14:09, Moll, Ralf a écrit : > Hi Denis, Hi Ralf, thank you for this feedback, [...] > > Here you can see that DAR triggers cp to copy the last slice again, > even though it is already present. This confirms my assumption. Note the slightly shift of interpretation you do: I would not say: "DAR triggers cp to copy the last slice again, even though it is already present": But rather that Dar reads again the last slice (something which is known and expected). As such Dar invokes the -E provided command(s) with the slice number which will very shortly be need. Dar does not know about slice existence when it comes to run a -E command, precisely because it occurs *before* this time, for user to do what's necessary for the slice to be present in the expected directory (originally by loading a CD or DVD in a tray, but in fact anything you want or need like downloading or here copying the needed slice). It is thus, as you have seen, the -E command that should avoid to do anything unnecessary when the last slice is to be open again (for example using the %c context parameter provided by dar), like avoiding to download/copying it again if it has already been downloaded/copied. The use of -affs does not change this paradigm, but when an isolated catalog is provided, it avoids reading twice the same slice (and thus invoking the -E provided command(s) twice for the same slice number). > > Workarounds: > As you already suggested, either use cp -n to avoid overwriting > existing files, or keep the first slice very small using -S 1k. > > Cheers, > Cheers, Denis |
|
From: Moll, R. <me...@rm...> - 2026-04-21 13:17:30
|
Hi Denis,
as discussed, here are the results of my tests with different -E parameters.
== -E "cp .... " ==
touch temp/1234-2025_2025-08-14T07.16.50_.204.dar
dar -A 1234-2025_2025-08-14T07.16.50_CAT
-x temp/1234-2025_2025-08-14T07.16.50_
-R /media/raid10/restore
-g "1234-2025/25-0585/test.pdf"
-w
-K "$(cat 1234-2025_2025-08-14T07.16.50_password.txt)"
-E "cp /media/ltfs/%b.%N.%e %p"
Warning, the archive 0999316-2025_2025-08-14T07.16.50_ has been
encrypted. A wrong key is not possible to detect, it would cause DAR
to report the archive as corrupted
--------------------------------------------
7 inode(s) restored
including 0 hard link(s)
0 inode(s) not restored (not saved in archive)
0 inode(s) not restored (overwriting policy decision)
40 inode(s) ignored (excluded by filters)
0 inode(s) failed to restore (filesystem error)
0 inode(s) deleted
--------------------------------------------
Total number of inode(s) considered: 47
--------------------------------------------
EA restored for 0 inode(s)
FSA restored for 0 inode(s)
--------------------------------------------
dar works because it detects the name structure by the created "dummy"
last slice an copies it again. Better than copying it first manual and
than by dar again.
== -E "cp -n ...." ==
touch temp/1234-2025_2025-08-14T07.16.50_.204.dar
dar -A 1234-2025_2025-08-14T07.16.50_CAT
-x temp/1234-2025_2025-08-14T07.16.50_
-R /media/raid10/restore
-g "1234-2025/25-0585/test.pdf"
-w
-K "$(cat 1234-2025_2025-08-14T07.16.50_password.txt)"
-E "cp -n /media/ltfs/%b.%N.%e %p"
234-2025_2025-08-14T07.16.50_.204.dar has a bad or corrupted header,
please provide the correct file. [return = YES | Esc = NO]
Escaping...
Final memory cleanup...
Aborting program. User refused to continue while asking:
234-2025_2025-08-14T07.16.50_.204.dar has a bad or corrupted header,
please provide the correct file.
My interpretation: DAR attempts to copy the last slice again.
Normally, cp would overwrite the existing file, but due to the -n
parameter, overwriting is prevented. As a result, the restore fails,
if using a dummy file. Would work as aspected using the original last
slice.
== -E "echo cp ...." ==
touch temp/1234-2025_2025-08-14T07.16.50_.204.dar
dar -A 1234-2025_2025-08-14T07.16.50_CAT
-x temp/1234-2025_2025-08-14T07.16.50_
-R /media/raid10/restore
-g "1234-2025/25-0585/test.pdf"
-w
-K "$(cat 1234-2025_2025-08-14T07.16.50_password.txt)"
-E "echo cp /media/ltfs/%b.%N.%e %p"
cp /media/ltfs/1234-2025_2025-08-14T07.16.50_.204.dar /home/user01/temp/
234-2025_2025-08-14T07.16.50_.204.dar has a bad or corrupted header,
please provide the correct file. [return = YES | Esc = NO]
Escaping...
Final memory cleanup...
Aborting program. User refused to continue while asking:
234-2025_2025-08-14T07.16.50_.204.dar has a bad or corrupted header,
please provide the correct file.
Here you can see that DAR triggers cp to copy the last slice again,
even though it is already present. This confirms my assumption.
Workarounds:
As you already suggested, either use cp -n to avoid overwriting
existing files, or keep the first slice very small using -S 1k.
Cheers,
Ralf
Am Do., 16. Apr. 2026 um 21:03 Uhr schrieb Denis Corbin via
Dar-support <dar...@li...>:
>
>
> [...]
>
> >> == My questions ==
> >>
> >> Why are the required slices not automatically copied using the -E
> >> option? It seems that the slice numbering is not handled with three
> >> digits, even though .%N. is used.
> >
> > if having a doubt on the command you pass to dar, I would suggest using
> > dar's --empty option (dry-run) and prepend the command passed to -E by
> > an echo:
> > dar
> > [...]
> > --empty
> > -E "echo cp -n /media/ltfs/%b.%N.%e %p"
> >
>
>
> > this will show the command that should be executed without doing
> > anything, so you can quickly check and tune.
> Sorry this will not work unless you have slices available, even if
> nothing is restored nor modified, dar still needs to read the slice.
> You can however check with a small faking backup of many slices
> (smallest slice possible is a few hundred bytes...).
>
>
|
|
From: Denis C. <dar...@fr...> - 2026-04-16 21:05:02
|
Le 16/04/2026 à 22:08, Moll, Ralf a écrit : > Hi Denis, > > thank you very much for your quick and detailed reply — it really > helped me a lot. > > You mentioned that metadata is always present in both the first and > the last slice, and that during restore DAR will, by default, use the > last slice for metadata access. > > My question is: > should I already use the parameter during archive creation, for example: > > # Create DAR archive optimized for LTFS tape usage > dar -c backup \ > -R /data \ # Root directory to archive > -s 50G \ # Size of regular slices (data slices) > -S 1k \ # Size of the first slice (metadata anchor) > -@ catalog \ # Store isolated catalog > --alter=force-first-slice # Force DAR to use metadata from first slice > > or is --alter=force-first-slice only intended to be used during restore? this is only to be used (as an option) when reading an archive, thus when restoring, testing, diffing, listing... an archive. > > Your explanations regarding copying the slices make perfect sense to > me. Unfortunately, I will only be able to test this again on my system > next week. I plan to adapt my script and also evaluate the behavior of > cp, cp -n, and cp --update=all / cp --update=older in combination with > the placeholder file. > > As expected, extraction did not work using only the placeholder file > and the catalog. However, I assume that DAR may have used the > placeholder file to infer information such as the number of slices and > the numbering scheme, then fetched the last slice and subsequently the > required slice. That’s just a hypothesis for now — I will verify this > next week. this should be checked... but if you generate a placeholder file to drive dar to auto-detect the min-digit parameters, better directly set it as argument to dar, no? > > It’s really impressive to see how much thought has gone into > parameters like -s and -S, and to better understand their purpose. > > Separately, I was wondering whether it might be worth considering > having all required metadata available entirely within the catalog > file. If the archive structure or slice size changes, access could > still fall back to the first or last slice, but in most real-world > cases, the archive layout will likely remain unchanged. the metadata that leads dar to read the first or last slice the archive format version, which defines which field are present, options available in an archive, the archive table of content (files, attributes, data location in the archive are loaded from the isolated catalog). This is under consideration to store this information withing an isolated catalog in the next major release. Still need to check its feasibility... > > Wishing you and all DAR users a great Friday and a good start into the weekend. Thanks! wishing the same to you too! > > Best regards, > Ralf > Cheers, Denis |
|
From: Moll, R. <me...@rm...> - 2026-04-16 20:38:15
|
Hi Denis,
thank you very much for your quick and detailed reply — it really
helped me a lot.
You mentioned that metadata is always present in both the first and
the last slice, and that during restore DAR will, by default, use the
last slice for metadata access.
My question is:
should I already use the parameter during archive creation, for example:
# Create DAR archive optimized for LTFS tape usage
dar -c backup \
-R /data \ # Root directory to archive
-s 50G \ # Size of regular slices (data slices)
-S 1k \ # Size of the first slice (metadata anchor)
-@ catalog \ # Store isolated catalog
--alter=force-first-slice # Force DAR to use metadata from first slice
or is --alter=force-first-slice only intended to be used during restore?
Your explanations regarding copying the slices make perfect sense to
me. Unfortunately, I will only be able to test this again on my system
next week. I plan to adapt my script and also evaluate the behavior of
cp, cp -n, and cp --update=all / cp --update=older in combination with
the placeholder file.
As expected, extraction did not work using only the placeholder file
and the catalog. However, I assume that DAR may have used the
placeholder file to infer information such as the number of slices and
the numbering scheme, then fetched the last slice and subsequently the
required slice. That’s just a hypothesis for now — I will verify this
next week.
It’s really impressive to see how much thought has gone into
parameters like -s and -S, and to better understand their purpose.
Separately, I was wondering whether it might be worth considering
having all required metadata available entirely within the catalog
file. If the archive structure or slice size changes, access could
still fall back to the first or last slice, but in most real-world
cases, the archive layout will likely remain unchanged.
Wishing you and all DAR users a great Friday and a good start into the weekend.
Best regards,
Ralf
Am Do., 16. Apr. 2026 um 18:48 Uhr schrieb Denis Corbin via
Dar-support <dar...@li...>:
>
> Le 16/04/2026 à 15:04, Moll, Ralf a écrit :
> > Hello Denis,
>
> Hello Ralf,
>
> >
> > first of all, I would like to thank you for your excellent tool.
>
> thanks! :)
>
> > I am
> > currently using DAR in an extended proof of concept in our HQ to
> > archive individual cases to LTFS-formatted LTO-8 tapes in a
> > data-protection-compliant way.
> >
> > = DAR Version =
> > dar version 2.7.8 on Debian 12
> >
> > = Backup Creation =
> >
> > BCK_PAR="--min-digits 3 -s 51200M --hash md5 -vt -vm -M"
> > dar ${BCK_PAR} -@ "${BCK_CAT}CAT" -c "${BCK_DST}" -R "${BCK_SRC_DIR}"
> > -g "${BCK_SRC_BASE}" -Kaes:${BCK_PAS}
> >
> > = Created Files =
> >
> > Slices:
> >
> > 1234-2025_2025-08-14T07.16.50_.001.dar
> > 1234-2025_2025-08-14T07.16.50_.001.dar.md5
> > 1234-2025_2025-08-14T07.16.50_.002.dar
> > 1234-2025_2025-08-14T07.16.50_.002.dar.md5
> > ...
> > 1234-2025_2025-08-14T07.16.50_.204.dar
> > 1234-2025_2025-08-14T07.16.50_.204.dar.md5
> >
> > Each slice has a size of 50 GiB.
> >
> > Catalog file:
> >
> > 1234-2025_2025-08-14T07.16.50_CAT
> >
> > Password file:
> >
> > 1234-2025_2025-08-14T07.16.50_password.txt
> >
> > = Selective Restore =
> >
> > == Determining the relevant slice(s) using the catalog file ==
> >
> > dar -l 1234-2025_2025-08-14T07.16.50_CAT -Tslice -g "1234-2025/25-0585/test.pdf"
> >
> > Output:
> >
> > Slice(s)|[Data ][D][ EA ][FSA][Compr][S]|Permission| Filename
> > --------+--------------------------------+----------+-----------------------------
> > [InRef][-] [---][ 0%][ ] drwxrwxrwx 1234-2025
> > [InRef][-] [---][ 3%][ ] drwxrwxrwx 1234-2025/25-0585
> > 166 [InRef][ ] [-L-][-----][X] -rwxrwxrwx 1234-2025/25-0585/test.pdf
> > -----
> > All displayed files have their data in slice range [166]
> >
> > Based on the catalog, the file appears to be fully stored in slice 166.
> >
> > == Automatic copying of required slices via -E fails ==
> >
> > dar -A 1234-2025_2025-08-14T07.16.50_CAT \
> > -x temp/1234-2025_2025-08-14T07.16.50_ \
> > -R /media/raid10/restore \
> > -g "1234-2025/25-0585/test.pdf" \
> > -w \
> > -K "$(cat 1234-2025_2025-08-14T07.16.50_password.txt)" \
> >
> >
> > This results in:
> >
> > cp: cannot stat '/media/ltfs/1234-2025_2025-08-14T07.16.50_.0.dar': No
> > such file or directory
> >
> > followed by:
> >
> > Error during user command line execution: execution of [ cp
> > /media/ltfs/1234-2025_2025-08-14T07.16.50_.0.dar
> > /media/raid5/restore/temp ] returned error code: 256
> >
> > If I explicitly add --min-digits 3, I instead get:
> >
> > Error during user command line execution: execution of [ cp
> > /media/ltfs/1234-2025_2025-08-14T07.16.50_.000.dar
> > /media/raid5/restore/temp ] returned error code: 256
> > > I have read the following note:
> >
> > “Note that dar will initially require slice number zero, meaning the
> > last slice of the backup...”
> >
> > However, I assumed that this behavior would not be required when using
> > a separate catalog file.
>
> it is because you can change the slicing of an archive even after having
> created an isolated catalog. Thus the last slice cannot be known, but no
> worries, you can drive for dar request the first slice (read below).
>
> And yes, the table of content is loaded from the isolated catalogue, but
> there is still the archive format version to read from the
> backup/archive (which may differ from the isolated catalog if different
> version of dar have been used between backup and isolation). This
> information is located both at the beginning of the first slice and a
> the end of the last slice (which is where it is fetched by default).
>
> You should consider using the -affs option (--alter=force-first-slice)
> to fetch this information from the first slice, and use a tiny first
> slice ("-S 1k"(uppercase S) 1 KB is large enough for that). This way,
> you could store on disk only this tiny first slice and the isolated
> catalogue. Then using the -affs option, you will only load from tape the
> big slices requested to restore a particular file or set of files.
>
> Yes, there is room to improve things, this is something under
> consideration for version 2.9.0.
>
> In the meanwhile you can either store the first or the last slice out of
> tape, which ever is the smaller beside the catalog and only use -affs if
> the smaller is the first. For your next backup, you will surely create a
> tiny first slice...
>
> >
> > == Restore with only slice 166 and catalog still fails ==
> >
> > Even if I manually copy the required slice file (slice 166), the
> > restore still only works if the last slice is also present, although
> > the catalog file is provided:
> >
> > dar -A 1234-2025_2025-08-14T07.16.50_CAT \
> > -x temp/1234-2025_2025-08-14T07.16.50_ \
> > -R /media/raid10/restore \
> > -g "1234-2025/25-0585/test.pdf" \
> > -w \
> > -K "$(cat 1234-2025_2025-08-14T07.16.50_password.txt)"
> >
> > Output:
> >
> > Auto detecting min-digits to be 3
> > The last file of the set is not present ..., please provide it.
> >
> > == Additional observation regarding -E behavior ==
> >
> > I also observed that the command triggered by -E deletes a manually
> > provided last slice and then copies it again.
>
> if you mean the following argument you passed to dar:
> -E "cp /media/ltfs/%b.%N.%e %p"
>
> well... dar, by itself does not delete slices...
>
> But yes, this is the way the 'cp' unix command works. If you do not want
> this behavior, better using:
> -E "cp -n /media/ltfs/%b.%N.%e %p"
>
> (-n option of the unix cp command).
>
> >
> > It seems that the last slice is only required as a placeholder. In
> > fact, I can create it using:
> >
> > touch 1234-2025_2025-08-14T07.16.50_.204.dar
>
> this is not what I observe: I get this error message if I "touch" the
> last slice rather than copying it:
> backup.5.dar has a bad or corrupted header, please provide the correct
> file. [return = YES | Esc = NO]
>
> >
> > and the restore proceeds. This means that copying the full last slice
> > from tape could be avoided, saving time and bandwidth depending on
> > slice size.
>
> If so, replacing the -E option with the following script should solve
> your problem:
> -E "my_script.duc %p %b %N %e %c"
>
> > cat my_script.duc
> #!/bin/bash
>
> slicepath="/media/ltfs"
> slicename="$2.$3.$4"
> target="$1"
>
> if [ "$5" = "init" ] ; then
> touch "$target/$slicename"
> else
> cp -n "$slicepath/$slicename" "$target"
> fi
> >
>
> But I would be surprise if it was not issuing the warning I had while
> creating the last slice using "touch".
> [note that I have not tested this above script, it may fails upon syntax
> error, this is just here to expose the idea]
>
> >
> > Is this behavior expected or known?
>
> yes, so far.
>
> >
> > == My questions ==
> >
> > Why are the required slices not automatically copied using the -E
> > option? It seems that the slice numbering is not handled with three
> > digits, even though .%N. is used.
>
> if having a doubt on the command you pass to dar, I would suggest using
> dar's --empty option (dry-run) and prepend the command passed to -E by
> an echo:
> dar
> [...]
> --empty
> -E "echo cp -n /media/ltfs/%b.%N.%e %p"
>
> this will show the command that should be executed without doing
> anything, so you can quickly check and tune.
>
> > Is it mandatory to also specify
> > --min-digits 3, or is this behavior caused by DAR requiring the last
> > slice despite the presence of a catalog file?
>
> today, dar check the location where slices are to be read and determin
> the min-digit automatically. However, if there is no slice at all,
> obviously you need to provide the --min-digits argument for the -E
> option and its %N macro set the proper number of leader zeros...
>
> > Why does DAR require the last slice during extraction even when a
> > separate catalog file is provided with -A?
>
> should now have the answer from answer above, see also man page
> regarding -affs option to get more details.
>
> > This reduces robustness, as
> > the last slice must always be present and intact, even if the
> > requested file resides in a different slice.
>
> the last or the first slice.
>
> > Is the behavior of deleting and re-copying the last slice via -E
> > expected?
>
> -E is doing what the user tells it to run... nothing more, nothing less ;)
>
> > And is it intentional that a zero-length placeholder file is
> > sufficient for the restore process?
>
> it is not and it should unfortunately not work, check with the script I
> provided and let me know.
>
> >
> > Best regards,
> >
> > Ralf
> >
>
> Cheers,
> Denis
>
>
>
|
|
From: Denis C. <dar...@fr...> - 2026-04-16 19:02:57
|
[...] >> == My questions == >> >> Why are the required slices not automatically copied using the -E >> option? It seems that the slice numbering is not handled with three >> digits, even though .%N. is used. > > if having a doubt on the command you pass to dar, I would suggest using > dar's --empty option (dry-run) and prepend the command passed to -E by > an echo: > dar > [...] > --empty > -E "echo cp -n /media/ltfs/%b.%N.%e %p" > > this will show the command that should be executed without doing > anything, so you can quickly check and tune. Sorry this will not work unless you have slices available, even if nothing is restored nor modified, dar still needs to read the slice. You can however check with a small faking backup of many slices (smallest slice possible is a few hundred bytes...). |
|
From: Denis C. <dar...@fr...> - 2026-04-16 16:47:46
|
Le 16/04/2026 à 15:04, Moll, Ralf a écrit :
> Hello Denis,
Hello Ralf,
>
> first of all, I would like to thank you for your excellent tool.
thanks! :)
> I am
> currently using DAR in an extended proof of concept in our HQ to
> archive individual cases to LTFS-formatted LTO-8 tapes in a
> data-protection-compliant way.
>
> = DAR Version =
> dar version 2.7.8 on Debian 12
>
> = Backup Creation =
>
> BCK_PAR="--min-digits 3 -s 51200M --hash md5 -vt -vm -M"
> dar ${BCK_PAR} -@ "${BCK_CAT}CAT" -c "${BCK_DST}" -R "${BCK_SRC_DIR}"
> -g "${BCK_SRC_BASE}" -Kaes:${BCK_PAS}
>
> = Created Files =
>
> Slices:
>
> 1234-2025_2025-08-14T07.16.50_.001.dar
> 1234-2025_2025-08-14T07.16.50_.001.dar.md5
> 1234-2025_2025-08-14T07.16.50_.002.dar
> 1234-2025_2025-08-14T07.16.50_.002.dar.md5
> ...
> 1234-2025_2025-08-14T07.16.50_.204.dar
> 1234-2025_2025-08-14T07.16.50_.204.dar.md5
>
> Each slice has a size of 50 GiB.
>
> Catalog file:
>
> 1234-2025_2025-08-14T07.16.50_CAT
>
> Password file:
>
> 1234-2025_2025-08-14T07.16.50_password.txt
>
> = Selective Restore =
>
> == Determining the relevant slice(s) using the catalog file ==
>
> dar -l 1234-2025_2025-08-14T07.16.50_CAT -Tslice -g "1234-2025/25-0585/test.pdf"
>
> Output:
>
> Slice(s)|[Data ][D][ EA ][FSA][Compr][S]|Permission| Filename
> --------+--------------------------------+----------+-----------------------------
> [InRef][-] [---][ 0%][ ] drwxrwxrwx 1234-2025
> [InRef][-] [---][ 3%][ ] drwxrwxrwx 1234-2025/25-0585
> 166 [InRef][ ] [-L-][-----][X] -rwxrwxrwx 1234-2025/25-0585/test.pdf
> -----
> All displayed files have their data in slice range [166]
>
> Based on the catalog, the file appears to be fully stored in slice 166.
>
> == Automatic copying of required slices via -E fails ==
>
> dar -A 1234-2025_2025-08-14T07.16.50_CAT \
> -x temp/1234-2025_2025-08-14T07.16.50_ \
> -R /media/raid10/restore \
> -g "1234-2025/25-0585/test.pdf" \
> -w \
> -K "$(cat 1234-2025_2025-08-14T07.16.50_password.txt)" \
>
>
> This results in:
>
> cp: cannot stat '/media/ltfs/1234-2025_2025-08-14T07.16.50_.0.dar': No
> such file or directory
>
> followed by:
>
> Error during user command line execution: execution of [ cp
> /media/ltfs/1234-2025_2025-08-14T07.16.50_.0.dar
> /media/raid5/restore/temp ] returned error code: 256
>
> If I explicitly add --min-digits 3, I instead get:
>
> Error during user command line execution: execution of [ cp
> /media/ltfs/1234-2025_2025-08-14T07.16.50_.000.dar
> /media/raid5/restore/temp ] returned error code: 256
> > I have read the following note:
>
> “Note that dar will initially require slice number zero, meaning the
> last slice of the backup...”
>
> However, I assumed that this behavior would not be required when using
> a separate catalog file.
it is because you can change the slicing of an archive even after having
created an isolated catalog. Thus the last slice cannot be known, but no
worries, you can drive for dar request the first slice (read below).
And yes, the table of content is loaded from the isolated catalogue, but
there is still the archive format version to read from the
backup/archive (which may differ from the isolated catalog if different
version of dar have been used between backup and isolation). This
information is located both at the beginning of the first slice and a
the end of the last slice (which is where it is fetched by default).
You should consider using the -affs option (--alter=force-first-slice)
to fetch this information from the first slice, and use a tiny first
slice ("-S 1k"(uppercase S) 1 KB is large enough for that). This way,
you could store on disk only this tiny first slice and the isolated
catalogue. Then using the -affs option, you will only load from tape the
big slices requested to restore a particular file or set of files.
Yes, there is room to improve things, this is something under
consideration for version 2.9.0.
In the meanwhile you can either store the first or the last slice out of
tape, which ever is the smaller beside the catalog and only use -affs if
the smaller is the first. For your next backup, you will surely create a
tiny first slice...
>
> == Restore with only slice 166 and catalog still fails ==
>
> Even if I manually copy the required slice file (slice 166), the
> restore still only works if the last slice is also present, although
> the catalog file is provided:
>
> dar -A 1234-2025_2025-08-14T07.16.50_CAT \
> -x temp/1234-2025_2025-08-14T07.16.50_ \
> -R /media/raid10/restore \
> -g "1234-2025/25-0585/test.pdf" \
> -w \
> -K "$(cat 1234-2025_2025-08-14T07.16.50_password.txt)"
>
> Output:
>
> Auto detecting min-digits to be 3
> The last file of the set is not present ..., please provide it.
>
> == Additional observation regarding -E behavior ==
>
> I also observed that the command triggered by -E deletes a manually
> provided last slice and then copies it again.
if you mean the following argument you passed to dar:
-E "cp /media/ltfs/%b.%N.%e %p"
well... dar, by itself does not delete slices...
But yes, this is the way the 'cp' unix command works. If you do not want
this behavior, better using:
-E "cp -n /media/ltfs/%b.%N.%e %p"
(-n option of the unix cp command).
>
> It seems that the last slice is only required as a placeholder. In
> fact, I can create it using:
>
> touch 1234-2025_2025-08-14T07.16.50_.204.dar
this is not what I observe: I get this error message if I "touch" the
last slice rather than copying it:
backup.5.dar has a bad or corrupted header, please provide the correct
file. [return = YES | Esc = NO]
>
> and the restore proceeds. This means that copying the full last slice
> from tape could be avoided, saving time and bandwidth depending on
> slice size.
If so, replacing the -E option with the following script should solve
your problem:
-E "my_script.duc %p %b %N %e %c"
> cat my_script.duc
#!/bin/bash
slicepath="/media/ltfs"
slicename="$2.$3.$4"
target="$1"
if [ "$5" = "init" ] ; then
touch "$target/$slicename"
else
cp -n "$slicepath/$slicename" "$target"
fi
>
But I would be surprise if it was not issuing the warning I had while
creating the last slice using "touch".
[note that I have not tested this above script, it may fails upon syntax
error, this is just here to expose the idea]
>
> Is this behavior expected or known?
yes, so far.
>
> == My questions ==
>
> Why are the required slices not automatically copied using the -E
> option? It seems that the slice numbering is not handled with three
> digits, even though .%N. is used.
if having a doubt on the command you pass to dar, I would suggest using
dar's --empty option (dry-run) and prepend the command passed to -E by
an echo:
dar
[...]
--empty
-E "echo cp -n /media/ltfs/%b.%N.%e %p"
this will show the command that should be executed without doing
anything, so you can quickly check and tune.
> Is it mandatory to also specify
> --min-digits 3, or is this behavior caused by DAR requiring the last
> slice despite the presence of a catalog file?
today, dar check the location where slices are to be read and determin
the min-digit automatically. However, if there is no slice at all,
obviously you need to provide the --min-digits argument for the -E
option and its %N macro set the proper number of leader zeros...
> Why does DAR require the last slice during extraction even when a
> separate catalog file is provided with -A?
should now have the answer from answer above, see also man page
regarding -affs option to get more details.
> This reduces robustness, as
> the last slice must always be present and intact, even if the
> requested file resides in a different slice.
the last or the first slice.
> Is the behavior of deleting and re-copying the last slice via -E
> expected?
-E is doing what the user tells it to run... nothing more, nothing less ;)
> And is it intentional that a zero-length placeholder file is
> sufficient for the restore process?
it is not and it should unfortunately not work, check with the script I
provided and let me know.
>
> Best regards,
>
> Ralf
>
Cheers,
Denis
|
|
From: Moll, R. <me...@rm...> - 2026-04-16 13:35:26
|
Hello Denis,
first of all, I would like to thank you for your excellent tool. I am
currently using DAR in an extended proof of concept in our HQ to
archive individual cases to LTFS-formatted LTO-8 tapes in a
data-protection-compliant way.
= DAR Version =
dar version 2.7.8 on Debian 12
= Backup Creation =
BCK_PAR="--min-digits 3 -s 51200M --hash md5 -vt -vm -M"
dar ${BCK_PAR} -@ "${BCK_CAT}CAT" -c "${BCK_DST}" -R "${BCK_SRC_DIR}"
-g "${BCK_SRC_BASE}" -Kaes:${BCK_PAS}
= Created Files =
Slices:
1234-2025_2025-08-14T07.16.50_.001.dar
1234-2025_2025-08-14T07.16.50_.001.dar.md5
1234-2025_2025-08-14T07.16.50_.002.dar
1234-2025_2025-08-14T07.16.50_.002.dar.md5
...
1234-2025_2025-08-14T07.16.50_.204.dar
1234-2025_2025-08-14T07.16.50_.204.dar.md5
Each slice has a size of 50 GiB.
Catalog file:
1234-2025_2025-08-14T07.16.50_CAT
Password file:
1234-2025_2025-08-14T07.16.50_password.txt
= Selective Restore =
== Determining the relevant slice(s) using the catalog file ==
dar -l 1234-2025_2025-08-14T07.16.50_CAT -Tslice -g "1234-2025/25-0585/test.pdf"
Output:
Slice(s)|[Data ][D][ EA ][FSA][Compr][S]|Permission| Filename
--------+--------------------------------+----------+-----------------------------
[InRef][-] [---][ 0%][ ] drwxrwxrwx 1234-2025
[InRef][-] [---][ 3%][ ] drwxrwxrwx 1234-2025/25-0585
166 [InRef][ ] [-L-][-----][X] -rwxrwxrwx 1234-2025/25-0585/test.pdf
-----
All displayed files have their data in slice range [166]
Based on the catalog, the file appears to be fully stored in slice 166.
== Automatic copying of required slices via -E fails ==
dar -A 1234-2025_2025-08-14T07.16.50_CAT \
-x temp/1234-2025_2025-08-14T07.16.50_ \
-R /media/raid10/restore \
-g "1234-2025/25-0585/test.pdf" \
-w \
-K "$(cat 1234-2025_2025-08-14T07.16.50_password.txt)" \
-E "cp /media/ltfs/%b.%N.%e %p"
This results in:
cp: cannot stat '/media/ltfs/1234-2025_2025-08-14T07.16.50_.0.dar': No
such file or directory
followed by:
Error during user command line execution: execution of [ cp
/media/ltfs/1234-2025_2025-08-14T07.16.50_.0.dar
/media/raid5/restore/temp ] returned error code: 256
If I explicitly add --min-digits 3, I instead get:
Error during user command line execution: execution of [ cp
/media/ltfs/1234-2025_2025-08-14T07.16.50_.000.dar
/media/raid5/restore/temp ] returned error code: 256
I have read the following note:
“Note that dar will initially require slice number zero, meaning the
last slice of the backup...”
However, I assumed that this behavior would not be required when using
a separate catalog file.
== Restore with only slice 166 and catalog still fails ==
Even if I manually copy the required slice file (slice 166), the
restore still only works if the last slice is also present, although
the catalog file is provided:
dar -A 1234-2025_2025-08-14T07.16.50_CAT \
-x temp/1234-2025_2025-08-14T07.16.50_ \
-R /media/raid10/restore \
-g "1234-2025/25-0585/test.pdf" \
-w \
-K "$(cat 1234-2025_2025-08-14T07.16.50_password.txt)"
Output:
Auto detecting min-digits to be 3
The last file of the set is not present ..., please provide it.
== Additional observation regarding -E behavior ==
I also observed that the command triggered by -E deletes a manually
provided last slice and then copies it again.
It seems that the last slice is only required as a placeholder. In
fact, I can create it using:
touch 1234-2025_2025-08-14T07.16.50_.204.dar
and the restore proceeds. This means that copying the full last slice
from tape could be avoided, saving time and bandwidth depending on
slice size.
Is this behavior expected or known?
== My questions ==
Why are the required slices not automatically copied using the -E
option? It seems that the slice numbering is not handled with three
digits, even though .%N. is used. Is it mandatory to also specify
--min-digits 3, or is this behavior caused by DAR requiring the last
slice despite the presence of a catalog file?
Why does DAR require the last slice during extraction even when a
separate catalog file is provided with -A? This reduces robustness, as
the last slice must always be present and intact, even if the
requested file resides in a different slice.
Is the behavior of deleting and re-copying the last slice via -E
expected? And is it intentional that a zero-length placeholder file is
sufficient for the restore process?
Best regards,
Ralf
|
|
From: Denis C. <dar...@fr...> - 2026-03-14 22:58:23
|
Hi Adam, OK, problem found, understood and solved. You should not have to change your backup, this was just a data processing/reading bug. You have the 2.8.4.RC1 release candidate available at: https://dar.edrusb.org/dar.linux.free.fr/Interim_releases/ currently building a dar_static version for x86_64 if that can help, it will be placed beside source package when available. non-regression tests will run for a big week then the 2.8.4 will be released, if no problem has been found in the meantime. Thank you for your feedback and script to reproduce the bug, this helped and saved me a lot of time! Cheers, Denis Le 10/03/2026 à 23:35, Denis Corbin via Dar-support a écrit : > Le 08/03/2026 à 03:29, Adam Watkins a écrit : >> Hello, > > Hello, > > my message sent on Sunday does not show in the mailing-list, I assume it > has been lost... > >> >> I believe I may have found an issue with DAR sequential streaming on >> pipes in version 2.8.3. >> >> Environment: >> OS: Bazzite >> dar version 2.8.3 >> libdar 7.0.2 >> compiled Feb 19 2026 >> Linux (x86_64) >> > > Thanks for this contextual info! > What is the archive header? [output of "dar -l <archive> -q"] > >> After generating an archive, a later test with dar -t works, however >> with dar --sequential-read -t -, it does not, generating the following >> error and several paths that do not exist in my archive, root dir >> items appearing below the first error and this directory keeps growing >> as the error occurs. >> >> Skipping backward is not possible on a pipe > > this is obviously a bug... > >> >> This is for a tape workflow, with dar_xform and mbuffer etc, but I >> isolated the problem to sequential-read from dar. >> >> dar ... -c - | dar --sequential-read -t - >> >> If needed, I can provide additional logs or test cases. > > Yes, it would worth knowing the few first lines of output when starts > the loop you describe including a couple of lines that succeeded the > test right before. > > Also, the same portion of the archive testing in non-sequential read > mode would help. > > You can send me that directly by email to keep this information as > private as possible if you want or you can share it here if you prefer. > >> >> Thanks, >> Adam >> >> > > Regards, > Denis > |
|
From: Denis C. <dar...@fr...> - 2026-03-10 22:36:17
|
Le 08/03/2026 à 03:29, Adam Watkins a écrit : > Hello, Hello, my message sent on Sunday does not show in the mailing-list, I assume it has been lost... > > I believe I may have found an issue with DAR sequential streaming on > pipes in version 2.8.3. > > Environment: > OS: Bazzite > dar version 2.8.3 > libdar 7.0.2 > compiled Feb 19 2026 > Linux (x86_64) > Thanks for this contextual info! What is the archive header? [output of "dar -l <archive> -q"] > After generating an archive, a later test with dar -t works, however > with dar --sequential-read -t -, it does not, generating the following > error and several paths that do not exist in my archive, root dir items > appearing below the first error and this directory keeps growing as the > error occurs. > > Skipping backward is not possible on a pipe this is obviously a bug... > > This is for a tape workflow, with dar_xform and mbuffer etc, but I > isolated the problem to sequential-read from dar. > > dar ... -c - | dar --sequential-read -t - > > If needed, I can provide additional logs or test cases. Yes, it would worth knowing the few first lines of output when starts the loop you describe including a couple of lines that succeeded the test right before. Also, the same portion of the archive testing in non-sequential read mode would help. You can send me that directly by email to keep this information as private as possible if you want or you can share it here if you prefer. > > Thanks, > Adam > > Regards, Denis |
|
From: Adam W. <acw...@gm...> - 2026-03-08 02:30:02
|
Hello, I believe I may have found an issue with DAR sequential streaming on pipes in version 2.8.3. Environment: OS: Bazzite dar version 2.8.3 libdar 7.0.2 compiled Feb 19 2026 Linux (x86_64) After generating an archive, a later test with dar -t works, however with dar --sequential-read -t -, it does not, generating the following error and several paths that do not exist in my archive, root dir items appearing below the first error and this directory keeps growing as the error occurs. Skipping backward is not possible on a pipe This is for a tape workflow, with dar_xform and mbuffer etc, but I isolated the problem to sequential-read from dar. dar ... -c - | dar --sequential-read -t - If needed, I can provide additional logs or test cases. Thanks, Adam |
|
From: Per J. <per...@pm...> - 2026-02-08 11:38:19
|
Hello,
Mar 29, 2025 I announced early v2 work around "v2-0.6.17".
I would like to provide a follow-up: `dar-backup` v2 has matured
significantly and is now in a stable state.
Since v2-0.6.17:
*
Catalog handling has been redesigned; FULL/DIFF/INCR chains are
resolved automatically and Point-In-Time restore is supported.
* Configurable Restore tests are executed automatically after backup
creation.
*
A full unit and integration test suite has been added, covering
backup, verification, restore, and cleanup workflows.
*
Per-backup PAR2 redundancy configuration is supported (ratio and
placement configurable per definition).
*
Logging and error handling have been hardened.
*
Documentation has been reorganized and expanded.
The tool is currently validated against `dar` 2.7.19.
I will begin testing against the 2.8.x series next.
Next planned area of work is integration with dar’s PKI-based encryption
features, allowing certificate-driven encryption (perhaps) per backup
definition.
Feedback is welcome, particularly from users running `dar` 2.8.x:
https://github.com/per2jensen/dar-backup<https://github.com/per2jensen/dar-backup>
Best regards,
Per Jensen
|