mirror of
https://github.com/jsvine/pdfplumber.git
synced 2026-08-29 16:40:24 +08:00
8.0 KiB
8.0 KiB
Change Log
All notable changes to this project will be documented in this file. The format is based on Keep a Changelog.
[0.5.23] — UNRELEASED
Added
- Add
utils.resolve(non-recursive .resolve_all) (7a90630) - Add
page.annotsandpage.hyperlinks, replacing non-functionalpage.annos, and mirroring pdfminer's language ("annot" vs. "anno"). (aa03961) - Add
page/pdf.to_jsonandpage/pdf.to_csv(cbc91c6) - Add
relative=True/Falseparameter to.cropand.within_bbox; those methods also now raise exceptions for invalid and out-of-page bounding boxes. (047ad34) [h/t @samkit-jain]
Changed
- Remove
pdfminer.from_pathandpdfminer.loadas deprecated; nowpdfminer.openis the canonical way to load a PDF. (00e789b) - Simplify the logic in "text" table-finding strategies; in edge cases, may result in changes to results. (
d224202)
Fixed
- Fix
.extract_words, which had been returning incorrect results whenhorizontal_ltr = False(d16aa13) - Fix
utils.resize_object, which had been failing in various permutations (d16aa13) - Fix
lines_stricttable-finding strategy, which a typo had prevented from being usable (f0c9b85) - Fix
utils.resolve_allto guard against two known sources of infinite recursion (cbc91c6)
Development Changes
- Rename default branch to "stable," to clarify its purpose
- Reformat code with psf/black (
1258e09) - Add code linting via psf/black and flake8 (
1258e09) - Switch from nosetests to pytest (
1ac16dd) - Switch from pipenv to standard requirements.txt + python -m venv (
48eaa51) - Add GitHub action for tests + codecov (
b148fd1) - Add Makefile for building development virtual environment and running tests (
4c69c58) - Add badges to README.md (
9e42dc3) - Add Trove classifiers for Python versions to setup.py (
6946e8d) - Add MANIFEST.in (
eafc15c) - Add GitHub issue templates (
c4156d6)
[0.5.22] — 2020-07-18
Changed
- Upgraded
pdfminer.sixrequirement to==20200517(cddbff7) [h/t @youngquan]
Added
- Add support for
non_stroking_colorattribute oncharobjects (0254da3) [h/t @idan-david]
[0.5.21] — 2020-05-27
Fixed
- Fix
Page.extract_table(...)to returnNoneinstead of crashing when no table is found (d64afa8) [h/t @stucka]
[0.5.20] — 2020-04-29
Fixed
- Fix
.get_page_imageto prefer paths over streams, when possible (ab957de) [h/t @ubmarco] - Local-fix pdfminer.six's
.resolve_allto handle tuples and simplify (85f422d)
Changed
- Remove support for Python 2 and Python <3.3
[0.5.19] — 2020-04-16
Changed
- Add
utils.decimalizeperformance improvement (830d117) [h/t @ubmarco]
Fixed
- Fix un-referenced method when using "text" table-finding strategy (
2a0c4a2c) - Add missing object type
rect_edgetoobj_to_edges()(0edc6bfa)
[0.5.18] — 2020-04-01
Changed
- Allow
rectandcurveobjects also to be passed to "explicit_..._lines" setting when table-finding. (And disallow other types of dicts to be passed.)
Fixed
- Fix
utils.extract_textbug introduced in prior version
[0.5.17] — 2020-04-01
Fixed
- Fix and simplify obj-in-bbox logic (see commit
25672961) - Improve/fix the way
utils.extract_texthandles vertical text (see commit8a5d858b) [h/t @dwalton76] - Have
Page.to_imageuse bytes stream instead of file path (Issue #124 / PR #179) [h/t @cheungpat] - Fix issue #176, in which
Page.extract_tablesdid not pass kwargs toTable.extract[h/t @jsfenfen]
[0.5.16] — 2020-01-12
Fixed
- Prevent custom LAParams from raising exception (Issue #168 / PR #169) [h/t @frascuchon]
- Add
sixas explicit dependency (for now)
[0.5.15] — 2020-01-05
Changed
- Upgrade
pdfminer.sixrequirement to==20200104 - Upgrade
pillowrequirement>=7.0.0 - Remove Python 2.7 and 3.4 from
toxtests
[0.5.14] — 2019-10-06
Fixed
- Fix sorting bug in
page.extract_table() - Fix support for password-protected PDFs (PR #138)
[0.5.13] — 2019-08-29
Fixed
- Fixed PDF object resolution for rotation (PR #136)
[0.5.12] — 2019-04-14
Added
cdecimalsupport for Python 2- Support for password-protected PDFs
[0.5.11] — 2018-11-13
Added
- Caching for
.decimalize()method
Changed
- Upgrade to
pdfminer.six==20181108 - Make whitespace checking more robust (PR #88)
Fixed
- Fix issue #75 (
.to_image()custom arguments) - Fix issue raised in PR #77 (PDFObjRef resolution), and general class of problems
- Fix issue #90, and general class of problems, by explicitly typecasting each kind of PDF Object
[0.5.10] — 2018-08-03
Fixed
- Fix bug in which, when calling get_page_image(...), the alpha channel could make the whole page black out.
[0.5.9] — 2018-07-10
Fixed
- Fix issue #67, in which bool-type metadata were handled incorrectly
[0.5.8] — 2018-03-06
Fixed
- Fix issue #53, in which non-decimalize-able (non_)stroking_color properties were raising errors.
[0.5.7] — 2018-01-20
Added
.travis.yml, but failing on.to_image()
Changed
- Move from defunct
pycryptotopycryptodome - Update
pdfminer.sixto20170720
[0.5.6] — 2017-11-21
Fixed
- Fix issue #41, in which PDF-object-referenced cropboxes/mediaboxes weren't being fully resolved.
[0.5.5] — 2017-05-10
Added
- Access to
__version__from main namespace
Fixed
- Fix issue #33, by checking
decode_text's argument type
[0.5.4] — 2017-04-27
Fixed
- Pin
pdfminer.sixto version20151013(for now), fixing incompatibility
[0.5.3] — 2017-02-27
Fixed
- Allow
import pdfplumbereven if ImageMagick not installed.
[0.5.2] — 2017-02-27
Added
- Access to
curvepoints. (E.g.,page.curves[0]["points"].) - Ability for
.draw_lineto drawcurvepoints.
Changed
- Disaggregated "min_words_vertical" (default: 3) and "min_words_horizontal" (default: 1), removing "text_word_threshold".
- Internally, made
utils.decimalizea bit more robust; now throws errors on non-decimalizable items. - Now explicitly ignoring some (obscure)
pdfminerobject attributes. - Raw input for
.draw_linefrom a bounding box to((x, y), (x, y)), for consistency withcurve["points"]and withPillow's underlying method.
Fixed
- Fixed typo bug when
.rect_edgesis called before.edges
[0.5.1] — 2017-02-26
Added
- Quick-draw
PageImagemethods:.draw_vline,.draw_vlines,.draw_hline, and.draw_hlines. - Boolean parameter
keep_blank_charsfor.extract_words(...)andTableFindersettings.
Changed
- Increased default
text_toleranceandintersection_toleranceTableFinder values from 1 to 3.
Fixed
- Properly handle conversion of PDFs with transparency to
pillowimages. - Properly handle
pandasDataFrames as inputs to multi-draw commands (e.g.,PageImage.draw_rects(...)).
[0.5.0] - 2017-02-25
Added
- Visual debugging features, via
Page.to_image(...)andPageImage. (Introduceswandandpillowas package requirements.) - More powerful options for extracting data from tables. See changes below.
Changed
- Entirely overhaul the table-extraction methods. Now based on Anssi Nurminen's master's thesis.
- Disentangle
.cropfrom.intersects_bboxand.within_bbox. - Change default
x_toleranceandy_tolerancefor word extraction from5to3
Fixed
- Fix bug stemming from non-decimalized page heights. [h/t @jsfenfen]
[0.4.6] - 2017-01-26
Added
- Provide access to
Page.page_number
Changed
- Use
.page_numberinstead of.page_idas primary identifier. [h/t @jsfenfen] - Change default
x_toleranceandy_tolerancefor word extraction from0to5
Fixed
- Provide proper support for rotated pages
[0.4.5] - 2016-12-09
Fixed
- Fix bug stemming from when metadata includes a PostScript literal. [h/t @boblannon]
[0.4.4] - Mistakenly skipped
Whoops.
[0.4.3] - 2016-04-12
Changed
- When extracting table cells, use chars' midpoints instead of top-points.
Fixed
- Fix find_gutters — should ignore
" "chars