mirror of
https://github.com/jsvine/pdfplumber.git
synced 2026-08-29 16:40:24 +08:00
3.7 KiB
3.7 KiB
Change Log
All notable changes to this project will be documented in this file. Currently goes back to v0.4.3.
The format is based on Keep a Changelog.
[0.5.10] — 2018-08-03
Fixed
- Fix bug in which, when calling get_page_image(...), the alpha channel could make the whole page black out.
[0.5.9] — 2018-07-10
Fixed
- Fix issue #67, in which bool-type metadata were handled incorrectly
[0.5.8] — 2018-03-06
Fixed
- Fix issue #53, in which non-decimalize-able (non_)stroking_color properties were raising errors.
[0.5.7] — 2018-01-20
Added
.travis.yml, but failing on.to_image()
Changed
- Move from defunct
pycryptotopycryptodome - Update
pdfminer.sixto20170720
[0.5.6] — 2017-11-21
Fixed
- Fix issue #41, in which PDF-object-referenced cropboxes/mediaboxes weren't being fully resolved.
[0.5.5] — 2017-05-10
Added
- Access to
__version__from main namespace
Fixed
- Fix issue #33, by checking
decode_text's argument type
[0.5.4] — 2017-04-27
Fixed
- Pin
pdfminer.sixto version20151013(for now), fixing incompatibility
[0.5.3] — 2017-02-27
Fixed
- Allow
import pdfplumbereven if ImageMagick not installed.
[0.5.2] — 2017-02-27
Added
- Access to
curvepoints. (E.g.,page.curves[0]["points"].) - Ability for
.draw_lineto drawcurvepoints.
Changed
- Disaggregated "min_words_vertical" (default: 3) and "min_words_horizontal" (default: 1), removing "text_word_threshold".
- Internally, made
utils.decimalizea bit more robust; now throws errors on non-decimalizable items. - Now explicitly ignoring some (obscure)
pdfminerobject attributes. - Raw input for
.draw_linefrom a bounding box to((x, y), (x, y)), for consistency withcurve["points"]and withPillow's underlying method.
Fixed
- Fixed typo bug when
.rect_edgesis called before.edges
[0.5.1] — 2017-02-26
Added
- Quick-draw
PageImagemethods:.draw_vline,.draw_vlines,.draw_hline, and.draw_hlines. - Boolean parameter
keep_blank_charsfor.extract_words(...)andTableFindersettings.
Changed
- Increased default
text_toleranceandintersection_toleranceTableFinder values from 1 to 3.
Fixed
- Properly handle conversion of PDFs with transparency to
pillowimages. - Properly handle
pandasDataFrames as inputs to multi-draw commands (e.g.,PageImage.draw_rects(...)).
[0.5.0] - 2017-02-25
Added
- Visual debugging features, via
Page.to_image(...)andPageImage. (Introduceswandandpillowas package requirements.) - More powerful options for extracting data from tables. See changes below.
Changed
- Entirely overhaul the table-extraction methods. Now based on Anssi Nurminen's master's thesis.
- Disentangle
.cropfrom.intersects_bboxand.within_bbox. - Change default
x_toleranceandy_tolerancefor word extraction from5to3
Fixed
- Fix bug stemming from non-decimalized page heights. [h/t @jsfenfen]
[0.4.6] - 2017-01-26
Added
- Provide access to
Page.page_number
Changed
- Use
.page_numberinstead of.page_idas primary identifier. [h/t @jsfenfen] - Change default
x_toleranceandy_tolerancefor word extraction from0to5
Fixed
- Provide proper support for rotated pages
[0.4.5] - 2016-12-09
Fixed
- Fix bug stemming from when metadata includes a PostScript literal. [h/t @boblannon]
[0.4.4] - Mistakenly skipped
Whoops.
[0.4.3] - 2016-04-12
Changed
- When extracting table cells, use chars' midpoints instead of top-points.
Fixed
- Fix find_gutters — should ignore
" "chars