Remove pandas and bump to 0.0.2

This commit is contained in:
Jeremy Singer-Vine
2015-12-15 11:40:14 -05:00
parent 11874bb6b7
commit ee04595109
4 changed files with 83 additions and 13 deletions
+79 -8
View File
@@ -2,19 +2,20 @@ __WARNING: This software is in its very early days, might not work well, and mig
# PDFPlumber
Plumb a PDF for detailed information about each char, rectangle, line, et cetera.
Plumb a PDF for detailed information about each text character, rectangle, line, and image.
Built on [`pdfminer`](https://github.com/euske/pdfminer)/[`pdfminer.six`](https://github.com/goulu/pdfminer).
Built on [`pdfminer`](https://github.com/euske/pdfminer) and [`pdfminer.six`](https://github.com/goulu/pdfminer).
## Installation
```sh
pip install git+https://github.com/goulu/pdfminer#egg=pdfminer.six
pip install pdfplumber
```
## Usage
### Basic Example
```python
import pdfplumber
@@ -25,14 +26,84 @@ pdf = pdfplumber.from_path("path/to/file.pdf")
with open("path/to/file.pdf") as f:
pdf = pdfplumber.load(f)
print(pdf.chars)
print(pdf.rects)
print(pdf.lines)
if len(pdf.chars):
print(pdf.chars[0])
if len(pdf.rects):
print(pdf.rects[0])
if len(pdf.lines):
print(pdf.lines[0])
```
### Pandas Integration
### Loading a PDF
By default, `pdf.chars`, etc., will be a plain Python dictionary. But if you `pandas=True` to `pdfplumber.load`/`.from_file`, you'll receive those properties as [Pandas dataframes](http://pandas.pydata.org/pandas-docs/stable/dsintro.html#dataframe).
`pdfplumber` provides two main ways to load a PDF:
- `pdfplumber.load(file_like_object)`
- `pdfplumber.from_path("path/to/file.pdf")`
Both methods return an instance of the `pdfplumber.PDF` class.
### Objects
Each instance of `pdfplumber.PDF` provides access to six types of PDF objects. The following properties each return a Python list of the matching objects:
- `.chars`, each representing a single text character.
- `.annos`, each representing a single annotation-text character.
- `.lines`, each representing a single 1-dimensional line.
- `.rects`, each representing a single 2-dimensional rectangle.
- `.images`, each representing a single image.
- `.figures`, each representing a single figure.
### Object Properties
Each object is represented as a simple Python `dict`, with the following properties:
- `char` / `anno`:
- `pageid`: Page ID on which this character was found.
- `text`: E.g., "z", or "Z" or " ".
- `fontname`: Name of the character's font face.
- `size`: Font size.
- `adv`: Equal to text width * the font size * scaling factor.
- `upright`: Whether the character is upright.
- `height`: Height of the character.
- `width`: Width of the character.
- `x0`: Distance of left side of character from left side of page.
- `x1`: Distance of right side of character from left side of page.
- `y0`: Distance of bottom of character from bottom of page.
- `y1`: Distance of top of character from bottom of page.
- `top`: Distance of top of character from top of page.
- `doctop`: Distance of top of character from top of document.
- `kind`: "LTChar" / "LTAnno"
- `line`:
- `pageid`: Page ID on which this line was found.
- `height`: Height of line.
- `width`: Width of line.
- `x0`: Distance of left-side extremity from left side of page.
- `x1`: Distance of right-side extremity from left side of page.
- `y0`: Distance of bottom extremity from bottom of page.
- `y1`: Distance of top extremity bottom of page.
- `top`: Distance of top of line from top of page.
- `doctop`: Distance of top of line from top of document.
- `linewidth`: Thickness of line.
- `kind`: "LTLine"
- `rect`:
- `pageid`: Page ID on which this rectangle was found.
- `height`: Height of rectangle.
- `width`: Width of rectangle.
- `x0`: Distance of left side of rectangle from left side of page.
- `x1`: Distance of right side of rectangle from left side of page.
- `y0`: Distance of bottom of rectangle from bottom of page.
- `y1`: Distance of top of rectangle from bottom of page.
- `top`: Distance of top of rectangle from top of page.
- `doctop`: Distance of top of rectangle from top of document.
- `linewidth`: Thickness of line.
- `kind`: "LTRect"
- `image`: TK
- `figure`: TK
## Python Support
+1 -1
View File
@@ -3,7 +3,7 @@ import pdfminer
import pdfminer.pdftypes
pdfminer.pdftypes.STRICT = False
VERSION_TUPLE = (0, 0, 1)
VERSION_TUPLE = (0, 0, 2)
VERSION = ".".join(map(str, VERSION_TUPLE))
def load(file_or_buffer, **kwargs):
+3 -3
View File
@@ -5,13 +5,13 @@ import subprocess
base_reqs = [
"chardet",
"pycrypto",
"pdfminer.six"
"pdfminer.six>=20151013"
]
setup(
name="pdfplumber",
version="0.0.1",
version="0.0.2",
packages=find_packages(exclude=["test",]),
tests_require=[ "nose", "pandas" ] + base_reqs,
tests_require=[ "nose", "pandas>=0.17.1" ] + base_reqs,
install_requires=base_reqs,
)
-1
View File
@@ -4,5 +4,4 @@ envlist = py27,py31,py34
[testenv]
deps=nose
pandas
git+https://github.com/goulu/pdfminer#egg=pdfminer.six
commands=nosetests