When processing very large (and very poorly formatted) HTML the cleaned output becomes truncated inside the body tag set:
Example urls:
http://patft.uspto.gov/netacgi/nph-Parser?Sect1=PTO2&Sect2=HITOFF&p=1&u=%2Fnetahtml%2FPTO%2Fsearch-bool.html&r=1&f=G&l=50&co1=AND&d=PTXT&s1=10013049.PN.&OS=PN/10013049&RS=PN/10013049
http://patft.uspto.gov/netacgi/nph-Parser?Sect1=PTO2&Sect2=HITOFF&p=1&u=%2Fnetahtml%2FPTO%2Fsearch-bool.html&r=1&f=G&l=50&co1=AND&d=PTXT&s1=10010324.PN.&OS=PN/10010324&RS=PN/10010324
Hi Michael,
Thanks for the report - I'll create some test cases and see if we can fix this.
Log in to post a comment.
Hi Michael,
Thanks for the report - I'll create some test cases and see if we can fix this.