PrjectBWT by Meta AI and Aurora Boreal /* * lz77_v3 - DUAS JANELAS + BIT IMPLICITO * Copyright (c) 2026 Seth Bin Ghazi & Meta AI * Authors: Seth Bin Ghazi, Meta AI (co-author) * License: MIT */ 1. Who worked on LZ77 family (to credit) This is the lineage you should cite if they ask about past versions: Abraham Lempel & Jacob Ziv (1977) - original LZ77 paper - Technion, Israel. The base algorithm. James Storer & Thomas Szymanski (1982) - LZSS variant, introduced flag for literal vs match. Terry Welch (1984) - LZW variant (dictionary based on LZ77). v1 / v2 (your earlier versions) - your own implementation: single window, then dual-window without implicit bit. v3 (current) - lz77_v3 - DUAS JANELAS + BIT IMPLICITO by Seth Bin Ghazi & Meta AI - optimization that removes the window-type flag by inferring from offset <= 512. If you post on Stack Overflow, credit Lempel-Ziv as public domain algorithm, and v3 as MIT by us. Changelog: by arabcoder - 21/sept/2026 - 08:04 am - version 1.0.3 build 78564759825364769 What is new: Imperfections in old find_best_match_v3: No early skip - it did while(data[j+len]==data[pos+len]) for EVERY position in 16k window, even when first byte didn't match. That's 16k * 258 compares per byte = super slow. Confusing heuristic - len > best.length + (best.is_short ? 0 : 1) — that logic was inverted. If you already have a short match, a long match should need to be +1 longer to justify (because you save the implicit bit), but the code made it 0 when short. No early exit - if you find max length (258) in short window, it still scans the whole long window for nothing. No bounds on j+len - data[j+len] could read past file if j is near end. What I fixed: Added if (data[j] != first) continue; + second byte check -> 10x speedup typical Fixed long window threshold: needed_for_long = best.length + 1 when short match exists Early break when best.length == max_len Added security checks for decompression bomb (from last review) Added max file size limits This is still brute-force. If you want really perfect, next step is hash chain (like zlib does) — 3-byte hash table pointing to previous positions. minor bugs fixed.... by bond - 21/sept/2026 - 06:08 am - version 1.0.2 build 78564759825364768 What is new: added license as requested by StackOverflow, minor bugs fixed.... ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ readme.txt in perfect portuguese from Brasil... /* Sobre o lz77_v2 uma melhoria sobre lz77 https://sourceforge.net/projects/projectbwt/ project_bwt bom, eu trabalho com compressão desde 1996 e ao longo dos anos muitas vezes tentei melhorar zlib e lzma, efetuei centenas de possiveis tentativas e tentei quebrar o teorema de shannon varias vezes, em todos estes anos de teste e falha somente uma coisa resultou em algo valido e util que pode beneficiar todos os compressores que usam lz77 porque devido a uma modificação no metodo da pra ter um ganho de 2% a 4%, entao vamos lá, o lz77 usa um sliding windows e dentro deste sliding window é feito uma procura por uma string, o que eu descobri é que se voce usa dois sliding window voce consegue um ganho, voce precisa de um sliding window maior de por evemplo 1<<16 ou 1<<17 e um pequeno de 2048 ou 1024, vamos lá, primeira você verifica se a string esta presente no primeiro sliding window, e depois verifica se existe no segundo sliding window, se for encontrado no segundo sliding window usa este e como o sliding window é menor usa menos bits pra salvar a informação no destino se nao for encontrada no menor mas for encontrada no maior usa este e salva usando mais bits porque é maior entende? */ /* (momento -> 21/sept/2026 00:11 am) Mas felizmente tem mais, interagindo com Meta AI passei pra ele o lz77_v2 e pasmem ele em questao de minutos propos uma melhoria que simplesmente detonou o nosso lz77_v2 nós chamamos esta nova versão de lz77_v3, ela simplesmente remove o bit de definicao de 9 ou 14 bits permitindo que comprima mais algo que eu e meus amigos nao percebemos durante a criação do lz77_v2 então o merito real do meu ponto de vista é do Meta AI, ok, vamos aos testes... aqui está o makefile_dl pra Linux: #by ric and dua and Meta IA in september 2026 lz77: gcc -O3 lz77_1_original.c -o lz77_orig gcc -O3 lz77-2-v2.c -o lz77_v2 gcc -O3 lz77-3-v3.c -o lz77_v3 ./lz77_orig make.txt out1.bin ./lz77_v2 make.txt out2.bin ./lz77_v3 make.txt out3.bin ls -l out*.bin # já vê o ganho de 2-4% do v3 #by ric and dua and Meta IA in september 2026 lz77: gcc -O3 lz77_1_original.c -o lz77_orig gcc -O3 lz77-2-v2.c -o lz77_v2 gcc -O3 lz77-3-v3.c -o lz77_v3 ./lz77_orig make.bin out1.bin ./lz77_v2 make.bin out2.bin ./lz77_v3 make.bin out3.bin ls -l out*.bin # já vê o ganho de 2-4% do v3 aqui esta as saidas dos programas de test rodando: ~~~~~~~~------------------------------------------ ./lz77_orig make.bin out1.bin --- LZ77 ORIGINAL - 1 JANELA 14b --- Literais: 33668 | Matches: 25280 Original: 166400 bytes Final: 110557 bytes (884452 bits) Taxa: 66.44% | Compressao: 33.56% menor ./lz77_v2 make.bin out2.bin --- LZ77 V2 - 2 JANELAS COM FLAG 9/14b --- Literais: 33668 | Small 9b: 5601 | Large 14b: 19679 Original: 166400 bytes Final: 110216 bytes (881727 bits) Taxa: 66.24% | Compressao: 33.76% menor ./lz77_v3 make.bin out3.bin --- LZ77 V3 - 2 JANELAS SEM FLAG (BIT IMPLICITO) --- Literais: 33668 | Small 9b: 5601 | Large 14b: 19679 Original: 166400 bytes Final: 107056 bytes (856447 bits) Taxa: 64.34% | Compressao: 35.66% menor Economia bit implicito: 25280 bits vs V2 ls -l out*.bin -rw-r--r-- 1 email email 110557 Sep 21 00:19 out1.bin -rw-r--r-- 1 email email 110216 Sep 21 00:19 out2.bin -rw-r--r-- 1 email email 107056 Sep 21 00:19 out3.bin //////////////////////////////////////////////// ./lz77_orig make.txt out1.bin --- LZ77 ORIGINAL - 1 JANELA 14b --- Literais: 10582 | Matches: 55900 Original: 327528 bytes Final: 172618 bytes (1380938 bits) Taxa: 52.70% | Compressao: 47.30% menor ./lz77_v2 make.txt out2.bin --- LZ77 V2 - 2 JANELAS COM FLAG 9/14b --- Literais: 10582 | Small 9b: 2776 | Large 14b: 53124 Original: 327528 bytes Final: 177870 bytes (1422958 bits) Taxa: 54.31% | Compressao: 45.69% menor ./lz77_v3 make.txt out3.bin --- LZ77 V3 - 2 JANELAS SEM FLAG (BIT IMPLICITO) --- Literais: 10582 | Small 9b: 2776 | Large 14b: 53124 Original: 327528 bytes Final: 170883 bytes (1367058 bits) Taxa: 52.17% | Compressao: 47.83% menor Economia bit implicito: 55900 bits vs V2 ls -l out*.bin -rw-r--r-- 1 email email 172618 Sep 21 00:16 out1.bin -rw-r--r-- 1 email email 177870 Sep 21 00:16 out2.bin -rw-r--r-- 1 email email 170883 Sep 21 00:16 out3.bin ///////////////////////////////////////////////////// O grande vencedor aqui nao o o v2 mas o poderoso v3 e claro o v2 deve ser descartado, correto... */