Simple PHP wrapper library for Keboola Storage API.
Library is available as composer package. To start using composer in your project follow these steps:
Install composer
curl -s http://getcomposer.org/installer | php
mv ./composer.phar ~/bin/composer # or /usr/local/bin/composerCreate composer.json file in your project root folder:
{
"require": {
"php" : ">=8.1",
"keboola/storage-api-client": "^14.0"
}
}Install package:
composer installAdd autoloader in your bootstrap script:
require 'vendor/autoload.php';Read more in Composer documentation.
Table write:
require 'vendor/autoload.php';
use Keboola\StorageApi\Client;
use Keboola\Csv\CsvFile;
$client = new Client([
'token' => 'YOUR_TOKEN',
'url' => 'https://connection.keboola.com'
]);
$csvFile = new CsvFile(__DIR__ . '/my.csv', ',', '"');
$client->writeTableAsync('in.c-main.my-table', $csvFile);Table export to file:
require 'vendor/autoload.php';
use Keboola\StorageApi\Client;
use Keboola\StorageApi\TableExporter;
$client = new Client([
'token' => 'YOUR_TOKEN',
'url' => 'https://connection.keboola.com'
]);
$exporter = new TableExporter($client);
$exporter->exportTable('in.c-main.my-table', './in.c-main.my-table.csv', []);File downloads (Client::downloadFile(), Client::downloadSlicedFile() and TableExporter) are
not bounded by object size on any provider — a download of any realistic size can finish, as long
as it keeps making progress. The three providers share one deliberate transfer policy:
AWS (S3ClientFactory) |
Azure (BlobClientFactory) |
GCP (GcsClientFactory) |
|
|---|---|---|---|
| Request deadline | 12 h liveness backstop | 12 h liveness backstop | 12 h liveness backstop |
| Stall detection | below 1 KB/s for 60 s | below 1 KB/s for 60 s | below 1 KB/s for 60 s |
| Connect timeout | 10 s | 10 s | 10 s |
| Retries | awsRetries, default Client::DEFAULT_RETRIES_COUNT (15) |
5, exponential (BlobStorageRetryMiddleware) |
3 (Google client default) |
| Writes to disk | directly (SaveAs becomes Guzzle's sink) |
via php://temp, then copied |
via php://temp, then copied (downloadToFile()) |
| Effective size ceiling | ~3.3 TB at 80 MB/s | ~3.3 TB at 80 MB/s | ~3.3 TB at 80 MB/s |
Notes:
- The deadlines are liveness backstops, not size caps. Stall detection alone cannot guarantee termination: it only fires below 1 KB/s and needs the whole 60 s window under the limit, so a link crawling just above that would otherwise run for months (40 GB at 1 KB/s is over a year). They are sized so no healthy transfer of any plausible export can reach them.
- A deadline bounds one attempt, not the whole download. The retry policies treat a timeout like any other failure, so the worst case for one call is (retries + 1) deadlines: 16 on AWS, 6 on Azure, 4 on GCP.
- Retries restart the whole object transfer from the first byte on AWS and Azure, so each retry pays
full egress. Keep
awsRetrieslow if you download very large files. On GCP an interruption that still carried a 2xx response resumes from the last fetched byte with aRangeheader (Rest::downloadObject()); any other failure restarts from the first byte. - Guzzle's
read_timeoutoption is honoured only by itsStreamHandler. Withext-curlinstalled all three download clients end up on the cURL handler, where the equivalent isCURLOPT_LOW_SPEED_LIMIT/CURLOPT_LOW_SPEED_TIME.ext-curlis not a hard requirement of this package, and without it Guzzle falls back to theStreamHandler; thereread_timeouttakes over as the stall detection, because the body is no longer requested as a stream, so the handler drains it itself and a stalled read raises. - Azure downloads go through
BlobClientFactory::createDownloadClient(). The Azure SDK requests blob bodies with Guzzle'sstreamoption, which routes them to theStreamHandler: there thecurloptions andconnect_timeoutare ignored,timeoutis a per-read socket timeout rather than a deadline, and a stalled transfer used to end the copy silently — reporting a truncated file as success. The download client clears that option in a middleware, so the body is read by the cURL handler Guzzle would pick anyway and a stall raises an exception thatBlobStorageRetryMiddlewareretries. - On Azure and GCP the body is buffered in
php://temp(memory up to 2 MB, then a temporary file in the system temp dir) before it is copied to the destination, so a download needs its size in free temp space on top of the destination. Uploads are unaffected. - The Azure upload client (
BlobClientFactory::createClientFromConnectionString()) keeps a 10 s connect timeout and a 120 s deadline per request, i.e. per 4 MiB block (ABSUploader::CHUNK_SIZE) or per whole blob for a small single-request upload. Uploads never request a streamed body, so there the deadline has always applied. - The GCP policy is passed per download call (
GcsClientFactory::downloadOptions()) rather than configured on theStorageClient, because client-level options never reach a download:Rest::downloadObject()always sets its ownrestOptions, andRequestWrapper::getRequestOptions()picks the per-requestrestOptionsover the client-level ones with??instead of merging them.
See LICENSE file.