os.walk() python: xml representation of a directory structure, recursion
Asked Answered
B

3

6

So I am trying to use os.walk() to generate an XML representation of a directory structure. I seem to be getting a ton of duplicates. It properly places directories within each other and files in the right place for the first portion of the xml file; however, after it does it correctly it then continues traversing incorrectly. I am not quite sure why....

Here is my code:

def dirToXML(self,directory):
        curdir = os.getcwd()
        os.chdir(directory)
        xmlOutput=""

        tree = os.walk(directory)
        for root, dirs, files in tree:
            pathName = string.split(directory, os.sep)
            xmlOutput+="<dir><name><![CDATA["+pathName.pop()+"]]></name>"
            if len(files)>0:
                xmlOutput+=self.fileToXML(files)
            for subdir in dirs:
                xmlOutput+=self.dirToXML(os.path.join(root,subdir))
            xmlOutput+="</dir>"

        os.chdir(curdir)
        return xmlOutput  

The fileToXML, simply parses out the list so no need to worry about that.

The Directory Structure is simply:

images/
images/testing.xml
images/structure.xml
images/Hellos
images/Goodbyes
images/Goodbyes/foo
images/Goodbyes/bar
images/Goodbyes/square

and the resulting xml file became:

<structure>
<dir>
<name>images</name>
  <files>
    <file>
      <name>structure.xml</name>
    </file>
    <file>
      <name>testing.xml</name>
    </file>
  </files>
  <dir>
    <name>Hellos</name>
  </dir>
  <dir>
    <name>Goodbyes</name>
    <dir>
      <name>foo</name>
    </dir>
    <dir>
      <name>bar</name>
    </dir>
    <dir>
      <name>square</name>
    </dir>
  </dir>
  <dir>
    <name>foo</name>
  </dir>
  <dir>
    <name>bar</name>
  </dir>
  <dir>
      <name>square</name>
    </dir>
  </dir>
  <dir>
    <name>Hellos</name>
  </dir>
  <dir>
    <name>Goodbyes</name>
    <dir>
      <name>foo</name>
    </dir>
    <dir>
      <name>bar</name>
    </dir>
    <dir>
      <name>square</name>
    </dir>
  </dir>
  <dir>
    <name>foo</name>
  </dir>
  <dir>
    <name>bar</name>
  </dir>
  <dir>
    <name>square</name>
  </dir>
</structure>

Any help would be much appreciated!

Barhorst answered 20/1, 2010 at 21:10 Comment(0)
L
9

I'd recommend against using os.walk(), since you have to do so much to massage its output. Instead, just use a recursive function that uses os.listdir(), os.path.join(), os.path.isdir(), etc.

import os
from xml.sax.saxutils import escape as xml_escape

def DirAsXML(path):
    result = '<dir>\n<name>%s</name>\n' % xml_escape(os.path.basename(path))
    dirs = []
    files = []
    for item in os.listdir(path):
        itempath = os.path.join(path, item)
        if os.path.isdir(itempath):
            dirs.append(item)
        elif os.path.isfile(itempath):
            files.append(item)
    if files:
        result += '  <files>\n' \
            + '\n'.join('    <file>\n      <name>%s</name>\n    </file>'
            % xml_escape(f) for f in files) + '\n  </files>\n'
    if dirs:
        for d in dirs:
            x = DirAsXML(os.path.join(path, d))
            result += '\n'.join('  ' + line for line in x.split('\n'))
    result += '</dir>'
    return result

if __name__ == '__main__':
    print '<structure>\n' + DirAsXML(os.getcwd()) + '\n</structure>'

Personally, I'd recommend a much less verbose XML schema, putting names in attributes and getting rid of the <files> group:

import os
from xml.sax.saxutils import quoteattr as xml_quoteattr

def DirAsLessXML(path):
    result = '<dir name=%s>\n' % xml_quoteattr(os.path.basename(path))
    for item in os.listdir(path):
        itempath = os.path.join(path, item)
        if os.path.isdir(itempath):
            result += '\n'.join('  ' + line for line in 
                DirAsLessXML(os.path.join(path, item)).split('\n'))
        elif os.path.isfile(itempath):
            result += '  <file name=%s />\n' % xml_quoteattr(item)
    result += '</dir>'
    return result

if __name__ == '__main__':
    print '<structure>\n' + DirAsLessXML(os.getcwd()) + '\n</structure>'

This gives an output like:

<structure>
<dir name="local">
  <dir name=".hg">
    <file name="00changelog.i" />
    <file name="branch" />
    <file name="branch.cache" />
    <file name="dirstate" />
    <file name="hgrc" />
    <file name="requires" />
    <dir name="store">
      <file name="00changelog.i" />

etc.

If os.walk() worked more like expat's callbacks, you'd have an easier time of it.

Lazulite answered 20/1, 2010 at 23:10 Comment(5)
... and you reached the same conclusion already. Not sure why I didn't get the "1 new answer" warning. -_-;Lazulite
Haha its ok... I will give it to you! Thanks for the help :)!Barhorst
Hi @MikeDeSimone, I'm intrigued by this use of the for loop: result += '\n'.join(' ' + line for line in DirAsLessXML(os.path.join(path, item)).split('\n')) I've searched for documentation + examples on this but could not find any, do you have any links?Zaria
@Zaria It’s not technically a “for loop”. It’s known in Python as a “comprehension” and this case a “generator”. Searching the Python docs for those terms should give you a lot. Basically, something of the form expression for loop-var in sequence [ if condition ] is a “generator,” which acts like a function with a yield statement (the basic form of a generator).Lazulite
So in your case, “DirAsLessXML(os.path.join(path, item)).split('\n'))” gets called, yielding a series of lines. Each of these lines is assigned to “line”, in order, and used to evaluate the expression “‘ ‘ + line”, and the results of these expressions are fed as a sequence into the join function. The overall effect is to add an indent to all lines output from the DirAsLessXML function.Lazulite
G
6

Remove the two lines:

        for subdir in dirs:
            xmlOutput+=self.dirToXML(os.path.join(root,subdir))

you are recursing into the subdirectories; but that's redundant, because os.walk recurses itself.

Geomorphology answered 20/1, 2010 at 21:20 Comment(1)
That doesn't really work for me... is there a way to list only the files and only the directories in python... I need the xml file to be more tree like.Barhorst
B
0

I was attempting to use os.walk, but I saw that it didn't work with the recursive tree structure that I wanted to create in xml. I modified my code as follows and it produce the result I need:

def dirToXML(self,directory):
        curdir = os.getcwd()
        os.chdir(directory)
        xmlOutput=""

        pathName = string.split(directory, os.sep)
        xmlOutput+="<dir><name><![CDATA["+pathName.pop()+"]]></name>"
        for item in os.listdir(directory):
            if os.path.isfile(os.path.join(directory, item)):
                xmlOutput+="<file><name><![CDATA["+item+"]]></name></file>"
            else :
                xmlOutput+=self.dirToXML(os.path.join(directory,item))
        xmlOutput+="</dir>"

        os.chdir(curdir)
        return xmlOutput    
Barhorst answered 20/1, 2010 at 22:16 Comment(0)

© 2022 - 2024 — McMap. All rights reserved.