Is there a way to escape a CDATA end token in xml?

M

10

146

I was wondering if there is any way to escape a CDATA end token (]]>) within a CDATA section in an xml document. Or, more generally, if there is some escape sequence for using within a CDATA (but if it exists, I guess it'd probably only make sense to escape begin or end tokens, anyway).

Basically, can you have a begin or end token embedded in a CDATA and tell the parser not to interpret it but to treat it as just another character sequence.

Probably, you should just refactor your xml structure or your code if you find yourself trying to do that, but even though I've been working with xml on a daily basis for the last 3 years or so and I have never had this problem, I was wondering if it was possible. Just out of curiosity.

Edit:

Other than using html encoding...

Messere answered 21/10, 2008 at 21:54 Comment(5)

First, i accept the answer as correct but note: Nothing precludes someone from encoding > as > within CData to ensure embedded ]]> will not be parsed as CDEnd. It simply means it's unexpected and that & must FIRST be encoded as & too so that the data can be properly decoded. Users of the document must know to decode this CData too. It's not unheard of since part of the purpose of CData is to contain content that a specific consumer understands how to handle. Such a CData just can't be expected to be interpreted properly by any generic consumer. – Bluepencil 16/5, 2011 at 14:54

@nix, CDATA just provides an explicit way to declare text node content such that language tokens within (other than ]]>) do not get parsed. It specifically does not expand entity references like > for this reason, so in a CDATA block, that just means those four characters, not '>'. To put it in perspective: in the xml spec, all text content is called "cdata", not just these sequences ("character data"). Also it's not about specific consuming agents. (Such a thing does exist though -- processing instructions (<?target instruction?>). – Hypogeous 11/10, 2015 at 5:37

(I should add, even if this sort of thing runs contrary to the original intent of the node, all is fair in the long & torturous battle with XML. I just feel it could be useful for readers to know that <![CDATA[]]> was not actually designed for that purpose.) – Hypogeous 11/10, 2015 at 5:48

@Hypogeous CDATA was designed to allow anything: they are used to escape blocks of text containing characters which would otherwise be recognized as markup That implies CDATA too since it is also markup. But, in fact, you don't need the double encoding I implied. ]]> is an acceptable means of encoding a CDEnd within a CDATA. – Bluepencil 11/10, 2015 at 19:14

True, you wouldn't need double encoding -- but you would still need the agent to have special knowledge, since the parser wouldn't parse > as >. That's what you mean though, I think? That you could replace them as you see fit, after parsing? – Hypogeous 11/10, 2015 at 19:41

D

153

You cannot escape a CDATA end sequence. Production rule 20 of the XML specification is quite clear:

[20]    CData      ::=      (Char* - (Char* ']]>' Char*))

EDIT: This product rule literally means "A CData section may contain anything you want BUT the sequence ']]>'. No exception.".

EDIT2: The same section also reads:

Within a CDATA section, only the CDEnd string is recognized as markup, so that left angle brackets and ampersands may occur in their literal form; they need not (and cannot) be escaped using "<" and "&". CDATA sections cannot nest.

In other words, it's not possible to use entity reference, markup or any other form of interpreted syntax. The only parsed text inside a CDATA section is ]]>, and it terminates the section.

Hence, it is not possible to escape ]]> within a CDATA section.

EDIT3: The same section also reads:

2.7 CDATA Sections

[Definition: CDATA sections may occur anywhere character data may occur; they are used to escape blocks of text containing characters which would otherwise be recognized as markup. CDATA sections begin with the string "<![CDATA[" and end with the string "]]>":]

Then there may be a CDATA section anywhere character data may occur, including multiple adjacent CDATA sections inplace of a single CDATA section. That allows it to be possible to split the ]]> token and put the two parts of it in adjacent CDATA sections.

ex:

<![CDATA[Certain tokens like ]]> can be difficult and <invalid>]]>

should be written as

<![CDATA[Certain tokens like ]]]]><![CDATA[> can be difficult and <valid>]]>

Dodecagon answered 21/10, 2008 at 22:31 Comment(17)

Indeed. Well, I'm not an academic type but as I said in the question, I'm just curious about this. To be honest, I'll just take your word on this, because I can barely make sense out of the syntax used for the rule. Thanks for your answer. – Messere 21/10, 2008 at 23:17

It reads like this: Char* (the set of all character sequences) - (except) Char* ']]>' Char* (the set of all character sequences that include the substring ']]>'). – Dodecagon 22/10, 2008 at 9:12

Thanks for the extra clarification. I'm accepting your answer as the one that better addresses the question I asked. (S. Lott's answer provides a work-around, which is fine, although it doesn't specifically deal with an actual escape char or sequence. – Messere 22/10, 2008 at 12:1

This is not an academic question. Think about an RSS feed of a blog post that contains a discussion about CDATA. – Insuppressible 12/7, 2011 at 15:5

I meant "academic" in the sense: "interesting to discuss, but without practical use". Generally, CDATA is not useful, it's just a way to serialize XML text, and it's semantically equivalent to escaping special chars using character entities < > and ". Characters entities is the simplest, most robust and most general solution, so use that instead of CDATA sections. If you use a proper XML library (instead of building XML out of strings) you don't even have to think about it. – Dodecagon 12/1, 2012 at 10:26

I just got bitten by this one because I am trying to encode some compressed Javascript into a <script> tag like: <script>/*<![CDATA[*/javascript goes here/*]]>*/</script> and my javascript includes just that sequence! I like the idea of splitting into multiple CDATA sections ... – Grand 23/3, 2012 at 1:11

If you were to add a CDATA code snippet in Sublime Text, it would require that you escape the ending sequence (configuration of Sublime is done almost exclusively through JSON and XML files). – Ichthyic 24/4, 2013 at 21:27

@NickT Instead of escaping the ending text in Sublime, you can do this: ]${1:Delete me then move along--required to escape CDATA end-tag}]>. Tools > New Snippet... annoys me, because it prints the snippet template into a new file. I don't want it a new file, so I just duplicated the blank snippet text itself into another snippet file...hence the need. – Parve 5/12, 2014 at 22:4

I experienced this in the real world. While reading the wikipedia dump and writing another xml file I encountered this on the page for the National Transportation Safety Board. It contained US$>100 million (2013) for the budget in the infobox. The source xml contained [[United States dollar|US$]]>100 million (2013) which was translated to [[United States dollar|US$]]>100 million (2013) by the reader and the writer opted to use CDATA to escape the text and failed. – Clemen 15/10, 2015 at 14:3

@Dodecagon re: it's just a way to serialize XML text or binary (unprintable) data. re: Characters entities is the simplest, most robust and most general solution for text that might confuse the XML parser, but if there are lots of them, it may be more space efficient to use CDATA. – Centuple 11/11, 2015 at 18:1

re: If you use a proper XML library and a proper library will have methods for adding CDATA (printable or unprintable) which will deal with the escape for you, if it needs to. Using a proper library is definitely the way to go. – Centuple 11/11, 2015 at 18:32

Re @jesse-chisholm: I am not sure what you are trying to say. CDATA might be more space efficient, but not in a way that should matter, since nobody should be transferring xml data that is not gzipped. After parsing, the memory usage should be the same. – Dodecagon 16/11, 2015 at 13:5

@ddaa: I was referring to the comment

Characters entities is the simplest, most robust and most general solution, so use that instead of CDATA sections. If you use a proper XML library (instead of building XML out of strings) you don't even have to think about it.

I was agreeing that using a proper library was better than building XML by hand, but disagreeing that entities are always the most robust, because if you have lots of them, then a CDATA is more efficient. Either way a proper library will handle it for you. And gzip makes the data binary which really needs CDATA. – Centuple 17/11, 2015 at 14:31

so, the answer is obvious: ]]> must be replaced with: ]]>]]><![CDATA[, in other words: close the current CDATA, type a "normal" ]]> but escaping the closing > and then open another CDATA. This would to the trick. – Bs 30/5, 2016 at 11:30

The answer is correct. CDATA sections do not escape content. I disagree whether this is academic though. If you are using XML format to store content in CDATA sections, then you can't store any XML content since it cannot tell the difference between content and markup. For this reason, the design of XML is broken. It fails the fundamental rule of parsing and delimiters: that you can embed delimiters in content with escaping. The design of CDATA breaks this rule. There are plenty other things wrong with XML as well, like how it's entitled to mess with whitespace in content. Use JSON. – Diadelphous 22/3, 2017 at 20:57

My point is that CDATA is useless in XML. It adds no expressiveness (everything you can do with CDATA you can do without it) and it provides an idiom that invites incorrect an fragile patterns: producing XML by string interpolation, and consuming XML without a proper parser. Therefore CDATA must be avoided. Therefore limitations in CDATA are "academic". – Dodecagon 24/3, 2017 at 6:0

Good answer, though I'd actually call escaping replacing ]]> with ]]]]><![CDATA[>, which, as you demonstrated, works. – Elena 19/11, 2019 at 16:7

G

180

You have to break your data into pieces to conceal the ]]>.

Here's the whole thing:

<![CDATA[]]]]><![CDATA[>]]>

The first <![CDATA[]]]]> has the ]]. The second <![CDATA[>]]> has the >.

Gambetta answered 21/10, 2008 at 22:27 Comment(0)

D

153