Hello,
@Test
public void entityRefTest() throws Exception {
final String html = "<html><body>"
+ "�"
+ "</body></html>";
loadPage(html);
}
java.lang.IllegalArgumentException
at java.lang.Character.toChars(Character.java:2583)
at org.cyberneko.html.HTMLScanner.appendChar(HTMLScanner.java:1675)
at org.cyberneko.html.HTMLScanner.scanEntityRef(HTMLScanner.java:1384)
at org.cyberneko.html.HTMLScanner$ContentScanner.scan(HTMLScanner.java:2049)
at org.cyberneko.html.HTMLScanner.scanDocument(HTMLScanner.java:918)
at org.cyberneko.html.HTMLConfiguration.parse(HTMLConfiguration.java:499)
at org.cyberneko.html.HTMLConfiguration.parse(HTMLConfiguration.java:452)
at org.apache.xerces.parsers.XMLParser.parse(Unknown Source)
Cheers,
Sebastian Cato
Diff:
Diff:
I appreciate the markup. The call to loadPage is however significant to the test case and should not be removed. To clarify, entityRefTest was written to be a part of one of the classes inheriting from SimpleWebTestCase in the test suite for HtmlUnit (I used WebClientTest), and loadPage refers to SimpleWebTestCase#loadPage(String).
There error is obvious, and it is due to non handling of "�" entities by NekoHtml.
I searched about any 6-digits hexadecimals for entity, but couldn't find none.
Can you provide any 6-digits value handled by real browsers as a visible character? Because otherwise, we will let NekoHtml team decide what to put (ignored or trimmed character).
Diff:
That sounds right to me, in the documentation for Character#toChars it says that IllegalArgumentException is thrown if the specified codePoint is not a valid Unicode code point, so I would say catching that exception straight away would be the thing to do.
Now fixed in SVN.
Thanks for the tiny test case (adapted version is now in build: MalformedHtmlTest.entityWithInvalidUTF16Code).